Testing and Guarantees

Testing methodology

Most Unicode text libraries rely on example-based testing: a developer writes a handful of input/output pairs, runs them in CI, and calls it done. Example-based tests verify the specific cases the developer thought of. They say nothing about the rest.

disarm combines three techniques that are uncommon in this space: compile-time data integrity assertions, exhaustive domain coverage, and stated invariant specifications. We are not aware of another transliteration or slugification library that publishes all three, though we haven't audited every library in every language.


What "exhaustively tested" means

Testing rigor is a spectrum between conventional tests and full formal verification (mathematical proofs of correctness). disarm operates at the strongest level achievable without nightly-only tools:

Level What it proves Who does this
Example-based tests Specific inputs produce expected outputs Everyone
Property-based tests Random inputs satisfy stated properties (statistical confidence) ~5% of open-source projects
Exhaustive domain tests Every element in a bounded domain satisfies stated properties (certainty) disarm
Compile-time assertions Data integrity invariants that fail the build if violated (zero runtime cost) disarm
Stated invariant specs Properties stated as specifications with verification method documented disarm
Bounded model checking Machine-checked proofs of absence of panics, overflow, UB Future (requires nightly Rust)

The gap between property-based testing and exhaustive testing is the difference between "we checked 1,000 random Hangul syllables" and "we checked all 11,172 Hangul syllables." The former gives statistical confidence. The latter gives certainty.


How the alternatives compare

Library Language Tests Exhaustive testing
Unidecode Python ~200 example tests None
text-unidecode Python ~50 example tests None
anyascii Multi Basic round-trip + snapshot None
python-slugify Python ~80 example tests None
awesome-slugify Python ~30 example tests None
confusable_homoglyphs Python ~20 example tests None
pathvalidate Python Example + parametrize None
unidecode (Rust) Rust ~10 example tests None
disarm Rust + Python 2,900+ tests Compile-time assertions, exhaustive domain, stated invariants

These libraries are mature and widely used. The test counts above are approximate (based on public repos at time of writing) and may not reflect internal or downstream test suites. The point is not that they are poorly tested — example-based testing is the norm — but that disarm's approach is different in kind.


The three layers of assurance

Layer 1: Compile-time data integrity assertions (build.rs)

Every time cargo build runs, the build script reads all transliteration TSV data files and asserts:

Assertion Scope Consequence if violated
All default BMP table values are pure ASCII 5,000+ mappings Build fails
All SMP table values are pure ASCII All supplementary mappings Build fails
All 22 language override tables contain only ASCII values de, ru, ja, fa, ... Build fails
All 20,924 Hanzi pinyin values are pure ASCII Full CJK block Build fails
Default BMP table has ≥ 5,000 entries Truncation detection Build fails
Hanzi pinyin table has ≥ 20,000 entries Truncation detection Build fails
Confusables table has ≥ 1,000 entries Truncation detection Build fails

Additionally, hangul.rs contains const assertions verifying that the Hangul decomposition algorithm constants match the Unicode specification: - JUNGSEONG_COUNT == 21, JONGSEONG_COUNT == 28 - Total syllable count = 19 × 21 × 28 = 11,172 - Compatibility jamo range = 51 entries exactly

These assertions execute at compile time, not in CI. A release artifact cannot exist if any assertion fails.

Layer 2: Exhaustive domain tests

These tests iterate over every element in a bounded Unicode domain. Unlike property-based tests (which sample randomly), exhaustive tests leave zero untested inputs within their domain.

Domain Size What is verified
All Hangul syllables (U+AC00–U+D7A3) 11,172 romanize_hangul() returns Some, output is ASCII, non-empty, decomposition indices in bounds, round-trip formula correct
All compatibility jamo (U+3131–U+3163) 51 lookup_compat_jamo() returns Some, output is ASCII
Full BMP, ErrorMode::Ignore (U+0080–U+FFFF) 63,488 transliterate_impl() produces ASCII-only output for every codepoint
Full BMP idempotence 63,488 f(f(ch)) == f(ch) for every codepoint
All CJK Unified Ideographs (U+4E00–U+9FFF) 20,992 Output is ASCII, unmapped count < 200
15 Indic script blocks ~2,000 codepoints Every consonant/vowel/virama in the block is correctly classified
Determinism 10 × 100 runs Same mixed-script input produces identical output 100 times
Mark stacks on the confusable fold (tests/mark_stacking.rs) 22,407 Every base the four confusable tables reach, crossed with every mark that composes with it, stacked 1–12 and 33 deep, under all three digit policies: a fixed point, and complete under numeric and tr39. See Run length

Total exhaustive coverage: ~159,000+ individually verified codepoints.

Layer 3: Stated invariant specifications

Seven properties are stated as specifications, each with a documented verification method:

ID Invariant Statement Verification
I1 ASCII Passthrough ∀s: s.is_ascii() → f(s) = s Exhaustive (all 128 ASCII) + Hypothesis 500
I2 ASCII Output ∀s: f(s, errors='ignore').is_ascii() Structural argument, checked in Lean (formal/lean/Transliterate) + exhaustive BMP per code point (Rust) + Hypothesis 1,000 incl. SMP
I3 Idempotence ∀s: f(f(s)) = f(s) Follows from I1 and I2 (Lean) + exhaustive BMP per code point (Rust) + Hypothesis 500
I4 No Exceptions ∀s ∈ UTF-8, |s| ≤ 10 MiB: f(s) does not throw Hypothesis 1,000 + explicit edge cases
I5 Deterministic ∀s, n>0: f(s) called n times → same result 100× repeat on 10 mixed-script inputs
I6 No Input Size Cap ∀s: f(s) accepts s whatever its length Boundary test: 12 MiB accepted (#80 removed the cap)
I7 Output Length Bounded ∀s: |f(s)| ≤ |s|_bytes × 5 + |s|_chars Exhaustive per code point (worst case U+337F, ratio 5) + Hypothesis 1,000

I1–I3 and I7 are stated for tones=False and no runtime registrations; the scope and the argument behind I2 are in Exhaustive Testing. Each invariant is a test class with a docstring stating the property. The verification method combines exhaustive enumeration (where the domain is bounded) with Hypothesis property-based testing (where it is not).

See formal-verification.md for the full specification document.


What exhaustive testing does NOT cover

Exhaustive testing is not formal verification. We are precise about the boundary:

Area Why not verified Mitigation
PHF hash correctness Trusted from phf_codegen crate Functional tests exercise every lookup path
Linguistic accuracy Transliteration correctness is empirical, not provable by testing alone Extensive corpus from native speakers; 83 language reference tests
Unicode version drift New Unicode versions add codepoints CI tracks Unicode version; unknown chars handled by ErrorMode
Memory safety / UB Requires Miri (nightly-only) unsafe_code = "forbid" in Cargo.toml — zero unsafe anywhere
Absence of panics Requires Kani bounded model checking (nightly-only) Property tests with 1,000+ random inputs; no panics in 2,900+ tests; cargo-fuzz over arbitrary bytes and text (below)
Run length beyond the tested shapes An exhaustive sweep fixes the shape of its inputs, and a bounded check proves nothing about inputs longer than its bound: the Lean check of the fold's convergence stopped at five characters, and the failure it was meant to rule out needs ten tests/mark_stacking.rs stacks marks past every pass cap in the code; cargo-fuzz grows and repeats input freely. See Run length

Future: When nightly Rust is available in CI, we plan to add Kani bounded model checking — a form of formal verification that would prove absence of panics and overflow in romanize_hangul, indic_char_role, and decomposition arithmetic — and Miri UB detection.


Run length: what the bounded checks could not see

The confusable fold iterated to a fixed point under a cap of eight passes (#434, 0.11.1), on the argument that "each pass removes at least one mark, so it converges in a couple of iterations". That argument bounds the passes by the number of marks, which the input chooses. C + U+0327 composes to Ç, which folds back to C, so every pass takes one cedilla: nine used up the cap, and from ten a release build returned a string that was still confusable and folded again on a second call. The presets, the pipeline and skeleton_key iterate the same fold in loops of their own under the same cap, and canonicalize_strict joined them when #862 (0.15.0) moved its mark cap after the fold. The nightly fuzz run found it on 2026-09-26, two nights after fuzzing was added (#1071).

Every layer above had a reason it could not, and none of them was bad luck:

Layer Why it could not see a stack of ten
Lean model of the fold Every symbol of the failure is in its alphabet (C, U+0327, U+04AA), but the convergence check covers every string up to length five: at most four cedillas, five passes. No amount of checking at that bound reaches ten
Lean differential test It checks that the model and the library agree, and they shared the cap. On the failing input they agree on the wrong answer
Library sweeps and exhaustive tests Every one crosses a base with one or two marks, so none can take more than three passes. "No swept input reached the cap" was true by construction
Property tests (proptest, Hypothesis) They draw characters independently. With U+0327 at 1.4% of draws, the likeliest generator here, a run of ten after a c comes up about once in 10^19 strings
The debug_assert on the cap It fires only in a debug build, and only on input nothing generated
cargo-fuzz libFuzzer inserts repeated bytes and copies input into itself, so runs are cheap for it. It found the bug

A second finding the same day had the opposite cause (#1072). canonicalize caps stacked marks before the fold, and the fold can move a mark to another class: ģ (a cedilla, below) folds to ġ (a dot, above). ģ and three marks above is four characters, inside the Presets model's exhaustive bound, but that model's alphabet had no fold that moves a mark, so the bound never mattered. With ģ and a third mark above added to the alphabet, the model's length-4 check fails on 12 words, every one ģ and three marks above.

What changed. The loops no longer stop at a cap (#1071), canonicalize caps again after the fold (#1072), and tests/mark_stacking.rs tests run length on purpose, on disarm::api, in PR CI:

  • the fold on every base the four tables reach (their sources, their values' characters, and each canonical prefix of a source's decomposition) crossed with every mark that composes with it, canonically or through a composition exclusion, stacked 1–12 and 33 deep, under all three policies;
  • every builder under every policy and every profile, on each stack whose composition is itself a fold source, and on each of those beside one to four marks of another class;
  • one stack of 20,000 per cycle, which must settle in the fold, canonicalize_strict and skeleton_key inside a time a pass per mark could not meet.

The cases come from the tables in src/tables/data/, not a list of known cycles, so a table change that adds one is covered the day it lands, and floors on the derived sets keep the test from passing by finding nothing. Run on the code before each fix, it fails: all three tests before #1071 (in a release build too, where the assertion is compiled out, from C and ten cedillas exactly), and the builders before #1072.

What it means for the other bounded claims.

  • A loop cap is a claim about the input. Either prove a constant bound, or finish the work some other way when the cap is reached; an argument whose bound grows with the input does not justify a constant.
  • A bounded check says nothing above its bound. Each one should say how long its shortest possible counterexample could be, and a result is evidence only if that is within the bound.
  • A sweep varies which characters appear. Run length is a separate dimension, and a sweep that fixes it at one or two cannot see a defect that needs ten.
  • Differential testing against a model finds where they differ, never a defect the model shares.
  • A model's alphabet needs a representative of every class the code branches on. "A fold that moves a mark to another combining class" was not one of the Presets model's classes. It is now (Finding 8 in formal/lean/Presets/README.md), with the fix #1072 made as a model variant that the bounded checks show is a fixed point.

Conventional testing (still comprehensive)

The exhaustive testing layers sit on top of a conventional test suite that is itself unusually thorough:

Test suite overview

Category Tests Coverage
Python (pytest) 2,268 All public API functions
Rust (#[test]) 635 Core algorithms, tables, edge cases
Exhaustive domain (Rust) 16 Full BMP, Hangul, CJK, Indic
Stated invariants (Python) 16 I1–I7 specifications
Property-based (Hypothesis) 500+ examples/property Full Unicode input space
Property-based (proptest) Rust-side invariants Normalization, roundtrips
Total 2,900+

Per-language reference tests

Each of the 83 built-in language profiles has dedicated tests verifying:

  • Known transliteration pairs — reference texts with expected output (e.g., "Москва" → "Moskva" for Russian, "Київ" → "Kyiv" for Ukrainian)
  • Language override behavior — lang="xx" produces different output from the default table where expected
  • ISO 9 and GOST interaction — scholarly modes override language-specific mappings correctly

Security invariant tests

tests/test_security_invariants.py uses Hypothesis to verify that canonicalize() enforces its security contracts on any input:

Invariant Guarantee
Bidi stripping All 13 bidi override/isolate characters removed
Zero-width stripping All 9 zero-width characters removed
Confusable neutralization No cross-script confusables in output
NFKC normalization Output always in NFKC form
Whitespace collapse No consecutive whitespace in output
Idempotency canonicalize(canonicalize(x)) == canonicalize(x)

Fuzzing, coverage and mutation testing

Three measurements of the suites above, all report-only: none is part of the required All checks passed status. How to run each is in docs/contributing/testing.md under Fuzzing, coverage and mutation testing.

Tool What it adds CI
cargo-fuzz (fuzz/) Arbitrary bytes through decode_to_utf8, and arbitrary text through ten surfaces, each checked against its documented properties fuzz.yml: 60 s per target on pull requests touching the core, 15 min nightly
cargo-llvm-cov Which lines and branches of the core the Rust suite executes coverage.yml: job summary and lcov artifact
cargo-mutants Whether the Rust suite notices when a line of a security-critical module changes mutants.yml: weekly, six modules

Baselines

Measured on 2026-09-23 on four cores, at the commit that added the tools. Re-measure rather than trust these once the code moves.

Fuzzing. Every target ran for at least 180 s from the committed seeds in the default build, which keeps debug assertions and overflow checks on, and seven of them also for 180 s optimized (nightly-2026-09-01, cargo-fuzz 0.13.2, AddressSanitizer). No run found a panic, an overflow, a sanitizer report or a timeout: every failure was a property. Throughput follows the work per input: presets runs every builder at least twice and managed about 360 inputs a second, decode_bytes calls the decoder thirteen times per input at about 720, and slugify about 4,000.

The runs found documented properties that did not hold: six in the first runs, two more on 2026-09-24, three on 2026-09-26 and one on 2026-10-02. Each is reproduced through the public API in tests/fuzz_findings.rs (and tests/test_fuzz_findings.py where the binding reaches it), and each target asserts the full property once its finding is resolved:

Surface Documented Reproduction Resolution
slugify numeric entities are decoded; allow_unicode gives one slug for both spellings (#477) A numeric entity that fails to decode is skipped together with up to 14 bytes of the ASCII after it: "Q&#A session" gives q, "Tom &#and Jerry" gives tom, and "issue &#12; fixed" (a control character, refused) gives issue. The skip stops at the first non-ASCII byte, so "&#a\u0301" gives "" while its NFC gives \u00e1. Fixed: &# with no digit after it is text, as in HTML, and an entity that names no allowed character is dropped without the text after it. A hex letter carrying a combining mark is not a digit, so both normal forms decode alike. "Q&#A session" gives q-a-session.
find_unmapped_confusables, find_confusables, find_untranslatable "its byte offset in the input string"; find_confusables: "the character as it appeared in the input" find_unmapped_confusables("\u04aa\u0327", Latin) reports U+0327 at offset 0, where U+04AA is; find_untranslatable("x\ufe0f") reports U+FE0F at offset 0, where x is; find_confusables("\u0456\u0308", Latin) reports U+0457, which the input does not contain. The locators walk composed clusters and report every character of one at the cluster's start. Fixed: each character is located at the input character it came from, and the one reported is the input's there: a mark that composes with nothing at its own offset, a decomposed homoglyph as its base with the composed character's fold as target.
find_untranslatable "exactly the set run would replace/ignore/preserve" transliterate("\U0001f240") is [?]ben[?] and find_untranslatable reports nothing: the NFKC brackets around the ideograph have no romanization. Fixed: a character counts as recovered only when its whole NFKC form transliterates, so U+1F240 is reported, and errors="strict" raises on it.
sanitize_filename "The result is a fixed point" With preserve_extension=False, "a" + ".*" * 9 gives a._, which sanitizes to a. Each pass strips one trailing ._, and the pass loop stops at eight (MAX_PASSES), which src/filename.rs admits can cost idempotence. Fixed: a pass repeats its trailing-separator and dot strips until neither removes anything, so the input settles in one call, and a debug build asserts the pass bound is not reached.
slugify with allow_unicode and separator="" a valid slug is unchanged "\u1100 \u1161" gives the two conjoining jamo, whose slug is U+AC00; "\U00016d67,\U00016d67" does the same with Kirat Rai. Joining the words puts two characters that compose side by side after composition has run. Fixed: with an empty separator the joined slug is composed again, so the first call returns U+AC00.
transliterate, invariant I7 output bytes at most five per input byte plus one per input character With tones=True, U+337F gives zhu sh\u00ec hu\u00ec sh\u00e8: 18 bytes for 3. I1-I3 were scoped to tones=False; I7 was not. Scoped: I7 is stated for tones=False, as I1-I3 are, and docs/formal-verification.md says why. It bounds the ASCII normalizer; toned pinyin is a display form whose vowels are two bytes.
catalog_key_with, search_key_with idempotent under every digit policy (the Presets model, #1024, #1029) Found on 2026-09-24: "\ufffd\ufffd\U00016d67\x16\U00016d67" keys as the two Kirat Rai vowel signs, and the key of that is U+16D68. The control is stripped after the last step that composes. Fixed: both builders end with an NFC pass, as sort_key and ml_normalize do. No fixture row moved; three were added.
slugify with allow_unicode none of the circled and squared Latin symbols in the slug (#1028) Found on 2026-09-24: a slug kept U+24B6 under a separator of NULs, /, U+24B6 and d. The U+24B6 was the separator's, inserted as given between two words. Not a library defect: the target checked the whole slug where the property is about the words. It now checks the words, and SlugConfig::separator says a separator is inserted as given.
normalize_confusables idempotent, and complete: is_confusable is false on the output (#522) Found on 2026-09-26: \u04aa followed by eight U+0327 among marks of another class. \u04aa folds to C, and C + U+0327 composes to Ç, which folds back to C, so each pass takes one cedilla. The loop stopped at eight passes (MAX_CONFUSABLE_PASSES): the input used them all, which tripped the loop's debug assertion. One cedilla more (C and ten, or \u04aa and nine) and a release build returned a string still confusable, which folded again on a second call. The loop had had its cap since #434 (0.11.1). c + U+0327 and i + U+0309 cycle the same way. The same input made canonicalize_strict return Ç and then C, and skeleton_key return ç and then c: the presets and the pipeline run their own fold loops under the same cap. Fixed: input still changing at the cap is finished span by span, and a pass that only shortens a run of one mark is applied as many times as it holds at once. All four loops do this, so the fold settles on any number of marks at the cost of a few passes, not one per mark.
canonicalize idempotent (#416, #835) Found on 2026-09-26, by the presets target while the row above was being checked: \u01e7 + U+0327 + U+0367 + three U+0327 + U+0303. The cap, three marks of one class on a base, runs before the confusable fold, and the fold turned \u0123 (a cedilla, below) into \u0121 (a dot, above), so the g carried four marks above and the next call cut one. Under tr39 and preserve the pre-fold folds before the cap, so only numeric failed. Fixed: the cap runs again after the fold, and touches only text it cuts. Its check returns early on text with no standalone mark, since no character decomposes to more than three, so canonicalize is cheaper than before.
slugify with allow_unicode no leading or trailing ZWJ/ZWNJ (#711) Found on 2026-09-26 by the slugify target, in a pull request's smoke run: "ab\u200d6" with separator="6" gave ab\u200d. The edges were trimmed of joiners before the trailing separator came off, and a separator the caller makes of word characters can match the end of a word, so taking it off exposed the joiner before it. The stopword filter, which splits on the separator again, and a truncation could end the slug on the same kind of spot. Fixed: the joiners are trimmed from both edges once the last step has run, and from the head UniqueSlugifier cuts again after its partial separator comes off. The shape properties (S2) stay stated for separators with no word character: such a separator still takes the word's own 6 with it.
transliterate, and through it search_key_with and catalog_key_with idempotent under every digit policy; an untranslatable character is replaced, dropped or kept as the error mode says Found on 2026-10-02 by the presets target, in a pull request's smoke run: search_key of \u0bdf\u0bc4\u0bc4 gave \u0bdf\u0bc4, and the key of that was \u0bdf. The Brahmic loop reads a role from the offset in the block, and both are unassigned code points at a consonant's and a vowel sign's offset. A "consonant" with no romanization opened the context, so the "sign" after it was erased, one per call, instead of reaching the error mode. Under errors="ignore" the sign also stripped the a of the text before it: "ba\u09a9\u093f" gave bi. Fixed: only a consonant whose romanization ends in its inherent a opens the context. Across every Brahmic consonant followed by every sign of its block, output moved only after the 78 consonant slots with no romanization (76 unassigned, and U+0C5C and U+0CDC, new in Unicode 17) and after ten characters the offsets call consonants whose romanization has no a (the Malayalam fractions, the Telugu and Kannada NAKAARA POLLU), and for those ten only before an unassigned sign.

Coverage. Rust, cargo +nightly-2026-09-01 llvm-cov --no-default-features --branch over the Tier-1 Rust suite (1,237 tests; doctests are not instrumented): 92.4% of lines (14,281 of 15,458), 85.1% of branches, 92.1% of functions under src/. The lowest modules, generated tables aside, are the Layer-2 wrappers that the Python suite exercises through the binding and the Rust suite mostly does not:

Module Lines Branches
src/api/presets.rs 56.0% -
src/api/safety.rs 66.7% -
src/lib.rs 71.0% -
src/normalize.rs 77.8% 66.7%
src/whitespace.rs 79.9% 77.8%
src/api/mod.rs 84.1% 45.0%
src/api/transliterate.rs 85.3% -
src/api/text.rs 85.6% -

Python, pytest --cov=disarm over the Tier-1 Python suite: 96% of the 1,289 statements of the disarm package, lowest _compat.py (88%) and _api.py (95%). That measures the Python wrappers only; the Rust they call is the table above.

Mutation. cargo-mutants 27.1.0 over two of the six modules, against the Rust suite:

Module Mutants Caught Missed Timeout Unviable
src/invisibles.rs 69 63 5 1 0
src/hostname.rs 52 30 20 0 2

121 mutants took 44 minutes with two jobs. The timeout is subdivision_flag_len returning Some(0), which never advances the scan: a hang the timeout catches, as it should. The misses are the useful part. Nothing in the Rust suite fails when is_invisible_in_hostname always returns false, when strip_invisibles does nothing, or when has_compat_form always returns false: the hostname screen's invisible and compatibility-form checks (#605, #709, #1019) are asserted only by the Python suite (tests/test_hn_compat_and_mapping.py), which cargo-mutants does not run. The same holds for strip_variation_selectors, which can return "xyzzy" unnoticed, for is_default_ignorable_format, and for the IPv6-literal parser. Every other binding calls this code through the Rust core, so each miss is a Rust test worth writing.

tests/hostname_and_invisibles.rs writes them, through the public API: one character of every invisible class and of both compatibility shapes through the hostname screen, the IPv6-literal boundaries (seven colons and eight, one zone ID and two, the characters a literal may hold), strip_variation_selectors, the default-ignorable formats, and a subdivision flag skipped whole. Re-measured on 2026-09-24 with the same tool and options:

Module Mutants Caught Missed Timeout Unviable
src/invisibles.rs 69 68 0 1 0
src/hostname.rs 52 48 2 0 2

The timeout is the same hang, caught. The two misses are equivalent mutants, which no test can kill: each turns one || of is_invisible_in_hostname into &&, removing the zero-width and tag classes (is_zero_width && is_tag) or the tag and variation-selector classes (is_tag && is_variation_selector) from the union. Every character of those three classes is also Default_Ignorable_Code_Point and none is a bidi control, so the union's last clause, added for the Lean model's Detection finding 8, still flags each one and the function is unchanged on every scalar.


CI matrix

Every pull request runs ci.yml, on Ubuntu:

Axis Values
Python 3.12 for the full test suite. smoke.yml installs the built wheel and sdist on 3.11, the declared floor (requires-python), and exercises them there
Rust checks cargo fmt --check, cargo clippy -D warnings, cargo test
Python checks pytest, ruff lint, mypy strict mode, doctest

The test suite does not run on macOS or Windows. publish.yml builds the wheels for Linux, macOS and Windows on those platforms at release.


Unicode table update process

When Unicode versions are updated:

  1. Dependency update — bump unicode-segmentation, unicode-normalization, and confusable table crates
  2. Rebuild tables — build.rs regenerates PHF lookup tables from TSV source data at compile time. Compile-time assertions verify the new data is well-formed.
  3. Exhaustive tests — the full BMP and CJK domain tests verify invariants hold across any new characters
  4. Property tests — Hypothesis tests verify invariants still hold across the new character space
  5. Reference text tests — existing per-language tests confirm no behavioral changes for known inputs