Precompiled Pipelines¶
Ready-to-use multi-step text processing pipelines. Each is a single compiled Rust function with no pipeline construction overhead at call time.
Renamed in 0.11 (#430)
Three presets were renamed to describe their mechanism rather than imply a safety outcome. The old names are deprecated aliases, behave identically, and are removed in 1.0:
| Old name | New name |
|---|---|
security_clean |
canonicalize |
display_clean |
strip_format |
normalize_user_input |
canonicalize_strict |
"1.0" here is the commercial-support milestone defined in
RELEASING.md, not the next release — per that policy disarm
expects to stay below 1.0 for a long time, so these aliases are not going away
imminently. See the Upgrading guide for the full rename history
across versions, including the is_safe_hostname boolean-polarity inversion.
canonicalize¶
canonicalize ¶
canonicalize(text: str, *, digit_policy: str = 'numeric') -> str
Canonicalize text for security-sensitive comparison.
For cleaning untrusted input before comparison, this is the entry point. It does not make text safe to emit; encode at the sink.
It is half of a spoof-resistance answer, not all of it (#882). This reports what
two strings collapse to; find_confusables reports what looks like something
else. Measured over confusable-bench.v1, the six key reducers together catch 72
of 120 malicious identifiers and find_confusables catches 66 — but either one
firing catches 108, at 0 false positives on the 20 benign controls for each alone
and for the pair. The reducers take the evasion class (42/54) that the detector
cannot see, and the detector takes the composability class (31/31) the reducers
cannot. A registry needs both questions asked.
Two steps introduce ASCII, not one (#719). The leading NFKC is the obvious
one; the confusable fold is the second and reaches characters NFKC leaves
alone. U+2236 RATIO has no decomposition at all and becomes :,
U+2044 FRACTION SLASH becomes /, U+2216 SET MINUS becomes \.
232 code points reach ASCII by the fold alone, 76 of them producing one of
: = % & ? # / \. A string that carried no delimiter can leave here carrying
one — encode at the sink.
inspect_anomalies reports it as confusable when the word also carries an
ASCII letter, which is #633's gate and what keeps Привет from firing. A
delimiter-only string such as ∶∶∶ folds to ::: and is not reported.
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC → strip bidi/format → strip invisible classes (#413) → strip_control → strip_zero_width → collapse_whitespace → drop repeated marks → cap combining marks at 3 (anti-zalgo,
429) → NFC → confusables and NFC to a fixed point → drop repeated marks → cap¶
combining marks again (the confusable fold is iterated with NFC so TR39 skeletoning is normalization-stable
and the preset is idempotent — #416/#434). PRESETS lists the steps exactly.
Collapses fullwidth bypasses, neutralizes homoglyph spoofing, strips dangerous bidi overrides and soft hyphens, then normalizes whitespace (collapsing runs, stripping control chars and zero-width injections).
Scoped to identifiers, not body text (#624). The confusable fold runs toward
Latin, so it rewrites non-Latin text that has a Latin lookalike — Arabic alef
becomes l, Hebrew yod becomes ', and Ελληνικά comes back as
Eλλnvikά. It also removes U+200C, which Persian orthography requires.
Nothing flags this first, because ordinary Persian is not an anomaly. Use it on
usernames, hostnames, filenames and log lines; see Limitations (docs/limitations.md)
before pointing it at a sentence.
The fold is not a romanization (#907). It runs without transliterating first, so a
Cyrillic word reaches the Latin table as shapes rather than as sounds: Москва comes
back Mockba, where search_key and catalog_key give the romanization
moskva. Mockba is not a spelling of anything. It is what the confusable table
does to letters that no transliteration step has already handled, and catalog_key's
own step list states the rule: transliterate first, so non-Latin scripts are romanized
before confusables.
Here that order is a decision rather than an omission, because the collapse is the
point: an attacker's Latin Mockba and a Cyrillic Москва are meant to meet, and
the romanizing surfaces cannot make them —
find_key_collisions(["Москва", "Mockba"], key="search_key") is empty. Pick by what
you need. normalize applies Unicode normalization and nothing else, so the script
survives — but it is not an identity function, and NFC still recomposes e followed
by U+0301 into é. The key builders give a romanization, and this gives the
collapse.
Warning
Canonicalizes Unicode for comparison; it is not an output
sanitizer and provides no XSS/HTML/SQL/injection protection. The NFKC
step maps fullwidth lookalikes to live ASCII metacharacters by design
(< → <), so the output may be more important to context-encode
on the way out, not less. Encode at the sink; never emit this result
into markup or a query unescaped.
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644, #733). Read the
Upgrade notes of any minor release before deploying it against stored
values. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> canonicalize("Ηello Ꮤorld") # Greek Η + Cherokee Ꮤ → Latin
'Hello World'
Note
digit_policy, and why the default cannot move (#885). "tr39" folds digit
variants onto the letters they imitate, and over confusable-bench.v1 the six key
builders reach 92 of 120 malicious rows under it against 72 by default.
It is still not the right default, and the corpus does not show why. TR39's digit
mappings cover every non-Latin numeral system, not just the styled Latin ones:
Arabic-Indic zero folds to ., one to l, five to o. So the Arabic year
٢٠٢٤ keys as ٢.٢٤ and the Persian ۱۴۰۳ as l۴.۳. The 20 benign
controls that measured "zero false-positive cost" contain no non-Latin digits at
all, so the population that pays is not in the sample.
"numeric" therefore stays the default and is a genuine no-op — passing it
gives output byte-identical to not passing it, so no stored key moves. Pass
"tr39" when your inputs are Latin identifiers and the extra reach is worth it.
Do not pass it to text that may carry Arabic, Persian, Indic or Thai numerals.
"preserve" (#648) keeps a non-Latin numeral in its own script, and since #896
that holds here: the fold on the raw text and the preset's own fold both run under
the policy, where the earlier pre-pass left the preset's fold at the default and
the setting did nothing (#949).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 137,955 single
characters reduce to "" here (487 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
There is no on_empty here: this returns text rather than a key. The
four key builders take one.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip bidi/format → strip invisibles (#413) → strip_control → strip_zero_width → collapse_whitespace → drop_repeated_marks → strip_zalgo (#429) → NFC → fixed point(confusables → NFC) → drop_repeated_marks
from disarm import canonicalize
assert canonicalize("ℝ𝕖𝕒𝕝 𝕥𝕖𝕩𝕥") == "Real text"
assert canonicalize("Ηello Ꮤorld") == "Hello World"
ml_normalize¶
ml_normalize ¶
ml_normalize(text: str, *, lang: str | None = None, emoji: str = 'cldr', fold_case: bool = True) -> str
ML/NLP text normalization pipeline.
resolve deletions → NFKC → emoji→text → [transliterate] → strip_accents
→ emoji→text → [fold_case] → strip_control → strip_zero_width → collapse_whitespace → NFC
Produces clean, accent-free text suitable for tokenizers, embeddings, and feature extraction. Emoji are expanded to their CLDR short-name descriptions.
"Emoji" means the Unicode property, not the CLDR table. The annotation data also
names 326 code points that carry neither Emoji nor Extended_Pictographic —
the curly quotes, the dashes, the currency signs, the math operators — and naming
those inserts words into ordinary prose: film’s came back as
film right apostrophe s, one token to four with the possessive gone. They pass
through unchanged since 0.15.0. demojize called directly still names them.
Case folding is on by default. Turn it off for a cased downstream model:
folding is destructive and cannot be undone later in the chain, and an uncased
evaluation harness cannot measure what it costs. This is the one preset where
folding is a side effect rather than the point — catalog_key,
search_key, and sort_key fold because a key has to collide, so
they have no such switch.
fold_case=False does not mean "preserve the input untouched": strip_accents
still runs, so José becomes Jose with the capital kept. Use
normalize_confusables when diacritics must survive as well.
Warning
Not a security preset. It assumes trusted input, and the name
describes a use case rather than an operation, which is the one way to
pick it by mistake. Bidi controls, private-use characters and homoglyphs
all pass straight through: a right-to-left override survives, and
Cyrillic аpple does not fold onto apple. For anything
user-supplied reach for canonicalize, or
get_pipeline("llm_guardrail") when the text is headed for a model.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> ml_normalize("Café RÉSUMÉ")
'cafe resume'
>>> ml_normalize("München", lang="de")
'muenchen'
>>> ml_normalize("José Martínez", fold_case=False)
'Jose Martinez'
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 2,735 single
characters reduce to "" here (2,735 excluding the Private Use
Area), and so do most strings built from them — but not all: each regional
indicator reduces to "" alone, and two together name a flag
(U+1F1FA U+1F1F8 is flag: united states). A caller keying a table on this
has all of them, plus "no value", competing for one slot.
There is no on_empty here: this returns text rather than a key. The
four key builders take one.
Pipeline steps¶
resolve_deletions → NFKC → emoji→text → [transliterate] → strip_accents → emoji→text → [fold_case] → strip_control → strip_zero_width → collapse_whitespace → NFC
from disarm import ml_normalize
assert ml_normalize("Café RÉSUMÉ") == "cafe resume"
assert ml_normalize("München", lang="de") == "muenchen"
assert ml_normalize("I ❤️ Python 🐍") == "i red heart python snake"
# fold_case=False drops the case fold for a cased downstream model (#559);
# every other stage — including strip_accents — still runs.
assert ml_normalize("José Martínez", fold_case=False) == "Jose Martinez"
catalog_key¶
catalog_key ¶
catalog_key(text: str, *, lang: str | None = None, strict_iso9: bool = False, digit_policy: str = 'numeric', on_empty: str | None = None) -> str
Library catalog key generation pipeline.
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC → strip_bidi → strip invisibles → fold_case → (transliterate → confusables → strip_accents, to a fixed point) → fold_case → strip_control → strip_zero_width → collapse_whitespace → NFC
Produces a canonical deduplication key for bibliographic titles.
Warning
The confusable fold runs after transliteration, so it rarely sees a
homoglyph. Anything that romanizes to something other than its
lookalike is consumed before the fold can act: Cherokee Ꮃ looks
like W and romanizes to la, so Ꮃorld keys as laorld.
Cyrillic and Greek are not the exception this warning used to claim (#735). They romanize like every other non-Latin script, and a romanization is a sound, not a shape — so a letter that looks like one Latin letter routinely keys as a different one::
раураl → raural not "paypal" (Cyrillic а, р, у)
аррlе → arrle not "apple"
В → v looks like B
Ѕ → dz looks like S
Η → i Greek Eta, looks like H
Measured over the Cyrillic and Greek letter blocks, 29 of 96 and 31 of
129 letters key off their visual target. The ones that do line up —
а/a, е/e, о/o — line up because the sound and the
shape happen to agree, which is a coincidence of those letters rather
than a property of the pipeline.
Private-use characters are stripped, with the other invisible classes (#805).
This builds a key; screen adversarial input separately, and use
normalize_confusables or is_confusable if what you need is the visual
question.
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644). Read the
Upgrade notes of any minor release before deploying it against stored
keys. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> catalog_key(" Café RÉSUMÉ ")
'cafe resume'
>>> catalog_key("ΩMEGA café")
'omega cafe'
Note
digit_policy is more than a digit setting on this builder. "numeric" is
the default and a genuine no-op. Under "tr39" or "preserve" the whole
confusable table runs on the raw text, before the key is built and before
transliteration consumes what it reads (#896) — not only its digit rows. So a
homoglyph the default romanizes is folded first: a Cyrillic spelling of
paypal keys as paypal rather than raural. "tr39" also destroys the numeric reading of Arabic, Persian, Indic and Thai
digits. "preserve" keeps a numeral through the fold, and this builder's
transliteration then romanizes it anyway — a key that maps every script to Latin
cannot keep one. See canonicalize for the measurements and the trade (#885).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 139,867 single
characters reduce to "" here (2,399 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
on_empty reserves a sentinel for that case — the fix
sanitize_filename already made with _ (#485). It applies only when
the input was non-empty, so absence keeps its own key.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip_bidi → strip invisibles → fold_case → fixed point(transliterate → confusables → strip_accents) → fold_case → strip_control → strip_zero_width → collapse_whitespace → NFC
from disarm import catalog_key
assert catalog_key(" Café RÉSUMÉ ") == "cafe resume"
assert catalog_key("Москва", lang="ru") == "moskva"
assert catalog_key("Москва", lang="auto") == "moskva"
assert catalog_key("Müller", lang="de") == "mueller"
strip_format¶
strip_format ¶
strip_format(text: str) -> str
Strip bidi/format and invisible-injection vectors from rendered content.
strip bidi/format → strip invisibles (#413, rendering policy) →
strip control → strip zero-width → collapse_whitespace
Lightweight cleanup for user-submitted content destined for rendering. Strips bidirectional overrides (which can visually reorder text to hide malicious content), soft hyphens, control characters, and zero-width injections, then collapses runs of whitespace to single spaces.
Warning
"Display-safe" means visual hygiene (no bidi reordering, no invisible
injections) — not markup-safe. This does no HTML escaping and does
not strip <, >, &. When rendering into HTML, still escape at
the template/output layer; disarm is not an XSS defense.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> strip_format("hello\x00world\u200b!")
'helloworld!'
>>> strip_format(" spaced out ")
'spaced out'
Pipeline steps¶
strip_bidi → strip invisibles (#413, rendering policy) → strip_control → strip_zero_width → collapse_whitespace
from disarm import strip_format
assert strip_format("hello\x00world\u200b!") == "helloworld!"
assert strip_format(" spaced out ") == "spaced out"
assert strip_format("admin\u202euser") == "adminuser"
search_key¶
search_key ¶
search_key(text: str, *, lang: str | None = None, digit_policy: str = 'numeric', on_empty: str | None = None) -> str
Search index key generation pipeline.
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC → strip_bidi → strip invisibles → fold_case → transliterate → strip_accents → fold_case → strip_control → strip_zero_width → collapse_whitespace → NFC
Produces a case-insensitive, accent-insensitive, script-insensitive
lookup key. Like catalog_key but without confusable
normalization — lighter and faster for search indexes.
Warning
Homoglyph collisions here are a side effect of transliteration, not a
confusable fold. There is no confusables step in this pipeline under the
default digit_policy (see the note on it below).
Cyrillic аpple and Greek gοogle do collide with their Latin
spellings, because those letters romanize to a and o — but a
lookalike that romanizes to something else does not. Cherokee Ꮃ
looks like W and romanizes to la, so Ꮃorld keys as
laorld and never meets world. Private-use characters are stripped,
with the other invisible classes (#805). This builds a key; screen adversarial
input separately with is_confusable or has_anomalies.
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644). Read the
Upgrade notes of any minor release before deploying it against stored
keys. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> search_key(" Café RÉSUMÉ ")
'cafe resume'
>>> search_key("Москва")
'moskva'
>>> search_key("Über allen Gipfeln")
'uber allen gipfeln'
Note
digit_policy is more than a digit setting on this builder. "numeric" is
the default and a genuine no-op. Under "tr39" or "preserve" the whole
confusable table runs on the raw text, before the key is built and before
transliteration consumes what it reads (#896) — not only its digit rows. So a
homoglyph the default romanizes is folded first: a Cyrillic spelling of
paypal keys as paypal rather than raural. The fold also rewrites
|, " and the backtick, which this builder leaves alone by default. The case
fold and transliteration then make sources the fold did not see, so under those
two policies the builder runs to a fixed point (Finding 2 of the Lean model in
formal/lean/Presets); under the default it runs once. "tr39" also destroys the numeric reading of Arabic, Persian, Indic and Thai
digits. "preserve" keeps a numeral through the fold, and this builder's
transliteration then romanizes it anyway — a key that maps every script to Latin
cannot keep one. See canonicalize for the measurements and the trade (#885).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 139,870 single
characters reduce to "" here (2,402 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
on_empty reserves a sentinel for that case — the fix
sanitize_filename already made with _ (#485). It applies only when
the input was non-empty, so absence keeps its own key.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip_bidi → strip invisibles → fold_case → transliterate → strip_accents → fold_case → strip_control → strip_zero_width → collapse_whitespace → NFC
Under digit_policy="tr39" or "preserve" the pre-fold is the whole confusable table, not only its digit rows, and the list runs until the key stops changing. See digit_policy on the key builders.
from disarm import search_key
assert search_key("Café RÉSUMÉ") == "cafe resume"
assert search_key("Москва", lang="ru") == "moskva"
assert search_key("ΩMEGA", lang="auto") == "omega"
skeleton_key¶
skeleton_key ¶
skeleton_key(text: str, *, digit_policy: str = 'numeric', on_empty: str | None = None) -> str
A spoof key: the TR39 skeleton plus the prototype classes disarm keeps apart.
Pipeline: resolve deletions → strip_bidi → strip invisibles → strip_control → strip_zero_width → NFKC → confusables → prototype fold → fixed-point(fold_case → confusables → NFKC) → collapse_whitespace
The confusable fold runs twice, and the second pass is not redundant. The
table's entry for a homoglyph is often on the lowercase form, so a capital
the first pass cannot match becomes matchable the moment case is folded:
Ω (U+2126 OHM SIGN) reaches the fold as Ω, folds to ω, and only
then to w. With a single pass skeleton_key("Ω") returned ω while
skeleton_key("ω") returned w — and a key that is not a fixed point is
not a key. A second pass rather than a reorder: the first has to see cased
text or the prototype fold has nothing to work with.
NFKC runs inside the loop too, because the case fold and the confusable fold
can each leave the text decomposed: U+0390 case-folds to three code points,
and U+00A5 + U+0300 folds to Y beside a grave it composes with. The strip
steps run first, ahead of NFKC: a character they remove that sits between a
base and its mark would otherwise block the composition, and the key would see
the bare base. Both were found by the Lean model in formal/lean/Confusables,
and before them the key was not always a fixed point.
TR39 puts I, l and 1 in one equivalence class and O/0 in
another. disarm's table stops short of both — every member of the capital-I
family folds to I and stops there — so paypaI survives every other
surface intact. This closes it.
Why a separate builder. The letter half costs six collision groups in the
235,976 entries of /usr/share/dict/words, and Ione/lone is the only
ordinary-word merge among them. That price holds only on cased text: after a
case fold, I ≡ l is i ≡ l, and the same class costs 264 groups of
ordinary vocabulary — boiling/bolling, doit/dolt, ail/all.
A factor of 44. No existing key builder runs a confusable fold before folding
case, and catalog_key cannot be reordered to (#419).
Not for display. The output is a key, and it is more destructive than any
preset that forwards text — the same split canonicalize and
canonicalize_strict already make.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> skeleton_key("paypaI") # the class catalog_key cannot reach
'paypal'
>>> skeleton_key("paypal") == skeleton_key("paypaI")
True
>>> skeleton_key("SKU-1O0") # digits kept apart by default
'sku-1o0'
>>> skeleton_key("SKU-1O0", digit_policy="tr39") # ...and merged on request
'sku-loo'
The digit half is destructive by design. Under "tr39" every one of
SKU-100, SKU-1O0, SKU-IOO and SKU-l00 is one key, as are
v1.0.1, vI.O.I and vl.o.l. For a spoof detector that is the point;
for a deduplication key over anything carrying a part number, a version or an
ISBN it destroys the field.
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 137,955 single
characters reduce to "" here (487 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
on_empty reserves a sentinel for that case — the fix
sanitize_filename already made with _ (#485). It applies only when
the input was non-empty, so absence keeps its own key.
Pipeline steps¶
resolve_deletions → strip_bidi → strip invisibles → strip_control → strip_zero_width → NFKC → confusables → **prototype fold** → fixed-point(fold_case → confusables → NFKC) → collapse_whitespace
The class the other builders cannot reach¶
TR39 puts I, l and 1 in one equivalence class and O/0 in another. disarm's table
stops short of both: every member of the capital-I family folds to I and stops there. So
paypaI survives every other surface intact.
from disarm import canonicalize, catalog_key, skeleton_key
canonicalize("paypaI") # 'paypaI' — unchanged
catalog_key("paypaI") == catalog_key("paypal") # False
skeleton_key("paypaI") == skeleton_key("paypal") # True
Why a separate builder, and not a flag¶
The letter half costs six collision groups in the 235,976 entries of
/usr/share/dict/words — i/l, ian/lan, io/lo, ione/lone, iowa/lowa,
iowan/lowan. Five are proper nouns; Ione/lone is the only ordinary-word merge.
That price holds only on cased text. After a case fold, I ≡ l is i ≡ l and the same
class costs 264 groups of ordinary vocabulary: boiling/bolling, doit/dolt,
silverer/sliverer, ail/all. A factor of 44.
No existing key builder runs a confusable fold before folding case. catalog_key folds
case at step 3 and reaches its confusable step at step 6, and the two cannot be swapped —
fold-before-transliterate is required for idempotency (#419). Hence a builder of its own.
digit_policy — the half you have to ask for¶
"numeric" (default) applies the letter half only. "tr39" adds 1 ≡ l and 0 ≡ O,
which is what an identifier skeleton wants and what a deduplication key must not have:
| kind | inputs that become one key under tr39 |
|---|---|
| part number | SKU-100, SKU-1O0, SKU-IOO, SKU-l00 |
| plate | B01, BOI, BOl, B0I |
| version | v1.0.1, vI.O.I, vl.o.l |
| address | Flat 10, Flat IO, Flat lO |
For a spoof detector that is the point. For a deduplication key over anything carrying a
part number, a version or an ISBN it destroys the field — which is why catalog_key, whose
docstring says "a canonical deduplication key for bibliographic titles", is the worst
available home for it rather than the best.
Not for display
The output is a key. It is more destructive than any preset that forwards text, in the
same way canonicalize_strict is more destructive than canonicalize: the more
aggressive rule lives in the entry point whose contract says so.
sort_key¶
sort_key ¶
sort_key(text: str, *, lang: str | None = None, digit_policy: str = 'numeric', on_empty: str | None = None) -> str
Sort key generation pipeline.
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC → strip_bidi → strip invisibles → fold_case → transliterate-non-Latin → fold_case → strip_control → strip_zero_width → collapse_whitespace → drop repeated marks → cap marks at 3 → NFC
A case-insensitive collation key that, unlike search_key,
preserves base accented characters rather than folding them away.
It keeps the accent so accented and unaccented forms stay distinct
("Über" folds to "über", not "uber") and the accent survives
for a locale-aware collator. Non-Latin scripts are still folded to a
consistent Latin form ("Война" → "voyna") so cross-script titles
interfile. This is the collation counterpart to search_key, which
folds accents away for exact-match lookup — the two are deliberately not
interchangeable for accented Latin input.
Note: the result is a normalized string, not a UCA collation-weight key, so
comparing keys with plain codepoint ordering will not interfile über
with ASCII u… words. Pass the key to a Unicode/locale collator when
linguistically-correct order matters; the value here is that the accent is
preserved for it rather than folded away.
Because Latin letters are preserved verbatim, lang only affects
transliteration of non-Latin runs; an accented Latin letter is never expanded
by a language profile here (e.g. sort_key("Über", lang="de") is
"über", whereas search_key("Über", lang="de") is "ueber").
Warning
No confusable fold under the default digit_policy. As with
search_key, any homoglyph collision is a side effect of transliteration:
Cherokee U+13B3 romanizes to la rather than folding onto the W it
resembles. Private-use characters are stripped, with the other invisible
classes (#805). This produces a collation key, not a screen.
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644). Read the
Upgrade notes of any minor release before deploying it against stored
keys. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> sort_key("Война и мир")
'voyna i mir'
>>> sort_key("Über allen Gipfeln")
'über allen gipfeln'
>>> sort_key(" Café ")
'café'
Note
digit_policy is more than a digit setting on this builder. "numeric" is
the default and a genuine no-op. Under "tr39" or "preserve" the whole
confusable table runs on the raw text, before the key is built and before
transliteration consumes what it reads (#896) — not only its digit rows. So a
homoglyph the default romanizes is folded first: a Cyrillic spelling of
paypal keys as paypal rather than raural. The fold also rewrites
|, " and the backtick, which this builder leaves alone by default. The case
fold and transliteration then make sources the fold did not see, so under those
two policies the builder runs to a fixed point (Finding 2 of the Lean model in
formal/lean/Presets); under the default it runs once. "tr39" also destroys the numeric reading of Arabic, Persian, Indic and Thai
digits. "preserve" keeps a numeral through the fold, and this builder's
transliteration then romanizes it anyway — a key that maps every script to Latin
cannot keep one. See canonicalize for the measurements and the trade (#885).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 138,404 single
characters reduce to "" here (936 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
on_empty reserves a sentinel for that case — the fix
sanitize_filename already made with _ (#485). It applies only when
the input was non-empty, so absence keeps its own key.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip_bidi → strip invisibles → fold_case → transliterate-non-Latin → fold_case → strip_control → strip_zero_width → collapse_whitespace → drop_repeated_marks → strip_zalgo → NFC
Like search_key, it runs to a fixed point under a non-default digit_policy.
Unlike search_key, sort_key preserves base accented characters so
accented and unaccented forms stay distinct and the accent survives for a
locale-aware collator. Non-Latin scripts are still folded to a consistent Latin
form; Latin letters (including accented ones) are kept verbatim, so lang only
affects non-Latin runs. (The key is a normalized string, not a UCA weight key —
pass it to a Unicode collator when linguistically-correct order matters.)
from disarm import search_key, sort_key
# accents preserved for ordering (contrast search_key, which folds them away)
assert sort_key("Über") == "über"
assert search_key("Über") == "uber"
# a language profile never expands an accented Latin letter in a sort key
assert sort_key("Über", lang="de") == "über"
# non-Latin scripts are still folded to Latin so titles interfile
assert sort_key("Война и мир", lang="ru") == "voyna i mir"
assert sort_key("Café") == "café"
canonicalize_strict¶
canonicalize_strict ¶
canonicalize_strict(text: str, *, digit_policy: str = 'numeric') -> str
Strict Unicode canonicalization of user input — not an injection defense.
Warning
This normalizes Unicode; it does not make text safe to emit into
HTML, JS, URLs, SQL, or shells. It performs no escaping and does not
strip <, >, & — <script>alert(1)</script> passes through
unchanged, and the NFKC step can surface ASCII metacharacters from
fullwidth lookalikes (<script> → <script>). This is not XSS
or injection protection: encode at the output sink (framework
auto-escaping, DOMPurify, parameterized queries). Run this before that
encoder, never instead of it. The name predates this clarification.
Scoped to identifiers, not body text (#624). The confusable fold runs toward
Latin, so it rewrites non-Latin text that has a Latin lookalike — Arabic alef
becomes l, Hebrew yod becomes ', and Ελληνικά comes back as
Eλλnvikά. It also removes U+200C, which Persian orthography requires.
Nothing flags this first, because ordinary Persian is not an anomaly. Use it on
usernames, hostnames, filenames and log lines; see Limitations (docs/limitations.md)
before pointing it at a sentence.
The fold is not a romanization (#907). Москва comes back Mockba, where the
key builders give moskva; the confusable table runs without transliterating first,
so a non-Latin word reaches it as shapes rather than as sounds. See canonicalize for
why that order is deliberate and what to use instead.
Runs no transliteration step while neutralizing Unicode-level attack vectors: zalgo stacking, homoglyph spoofing, bidi overrides, zero-width injections, and control characters. It used to say it "preserves the original script", which the paragraph above disproves: the confusable fold rewrites individual letters, and only the romanization step is absent (#907).
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC → strip_bidi →
strip_zero_width → strip_control → strip invisible classes (#413) → (confusables
and NFC, then the cross-script mark strip, to a fixed point) → drop repeated marks
→ strip_zalgo → collapse_whitespace → NFC (invisibles are stripped before
zalgo-capping so they cannot split combining-mark runs, the cap follows the
cross-script mark strip for the same reason (#862), and the terminal NFC
recomposes any base+mark left adjacent by a stripped invisible — keeping the
output idempotent, #416/#413)
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644, #733). Read the
Upgrade notes of any minor release before deploying it against stored
values. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> canonicalize_strict("Hello, world!")
'Hello, world!'
>>> canonicalize_strict("p\u0430ypal") # Cyrillic а → Latin a
'paypal'
>>> canonicalize_strict("admin\u202euser") # RLO stripped
'adminuser'
Note
digit_policy folds digit variants before the key is built, and the builder's
own confusable fold runs under the same policy (#896). "numeric" is the default
and a genuine no-op; "tr39" reaches more spoofs but destroys the numeric reading
of Arabic, Persian, Indic and Thai digits; "preserve" keeps a non-Latin numeral
in its own script (#949). See canonicalize for the measurements and the trade
(#885).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 137,955 single
characters reduce to "" here (487 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
There is no on_empty here: this returns text rather than a key. The
four key builders take one.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip_bidi → strip_zero_width → strip_control → strip invisibles (#413) → fixed point(fixed point(confusables → NFC) → strip_cross_script_marks) → drop_repeated_marks → strip_zalgo → collapse_whitespace → NFC
from disarm import canonicalize_strict
assert canonicalize_strict("Hello, world!") == "Hello, world!"
assert canonicalize_strict("p\u0430ypal") == "paypal"
assert canonicalize_strict("admin\u202euser") == "adminuser"
Unlike canonicalize, this pipeline also strips zalgo text (excessive combining mark stacking). Unlike catalog_key/search_key, it does not transliterate — the original script is preserved.
strip_obfuscation¶
strip_obfuscation ¶
strip_obfuscation(text: str, *, digit_policy: str = 'numeric') -> str
Maximum-strength text deobfuscation.
Neutralizes homoglyph spoofing, zalgo abuse, invisible character injection, and bidi attacks. Uses TR39 confusable mapping (visual similarity) — Cyrillic р→p, с→c, В→B — not phonetic transliteration.
Warning
Not an output sanitizer. Resolves Unicode obfuscation only; performs
no HTML/JS/SQL escaping and does not strip <, >, &. The NFKC
step folds fullwidth < to a live <, so the output can be more
important to encode than the input. Encode at the output sink — this is
not XSS or injection protection.
Does not transliterate. Non-Latin scripts that have no Latin
confusable equivalent pass through unchanged. Chain with
transliterate() explicitly if you also need romanization.
Read the exclusion in that sentence literally: the scripts that do have a
Latin confusable equivalent are rewritten. 22 Arabic code points, 12 Hebrew and
65 Greek fold to ASCII, so wholly non-Latin text comes back with Latin letters
in it. This preset also strips combining marks, which takes Indic vowel signs
along with Latin accents — বাংলা becomes বল, which is not a word.
Scoped to identifiers, not body text (#624). The confusable fold runs toward
Latin, so it rewrites non-Latin text that has a Latin lookalike — Arabic alef
becomes l, Hebrew yod becomes ', and Ελληνικά comes back as
Eλλnvikά. It also removes U+200C, which Persian orthography requires.
Nothing flags this first, because ordinary Persian is not an anomaly. Use it on
usernames, hostnames, filenames and log lines; see Limitations (docs/limitations.md)
before pointing it at a sentence.
The fold is not a romanization (#907). Москва comes back Mockba, where the
key builders give moskva; the confusable table runs without transliterating first,
so a non-Latin word reaches it as shapes rather than as sounds. See canonicalize for
why that order is deliberate and what to use instead.
Preserves case. Case is not deception — proper nouns, acronyms,
and sentence boundaries are meaningful. Chain with fold_case()
if lowercasing is also needed.
Pipeline: resolve deletions → [digit-policy pre-fold] → NFKC →
strip_zalgo(max_marks=0) → strip_bidi → strip_zero_width → strip invisibles →
confusables → strip_accents → strip_control → collapse_whitespace → NFC. There
is no demojize step (#910): an emoji is left where it stands, because a
comparison surface must not write attacker-chosen words into the value being
compared. The terminal NFC recomposes two characters a stripped control had kept
apart, keeping the output idempotent.
Note
Stability. A patch upgrade never changes this function's output; a
minor upgrade may, and is a possible reindex event (#644, #733). Read the
Upgrade notes of any minor release before deploying it against stored
values. The contract, and what has moved so far, is in docs/RUST_API.md
under Key stability.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> strip_obfuscation("P\u0430yP\u0430l") # Cyrillic а → Latin a
'PayPal'
>>> strip_obfuscation("\u0420rodu\u0441t") # Cyrillic Р→P, с→c
'Product'
>>> strip_obfuscation("H\u0338a\u0338t\u0338e\u0338 speech")
'Hate speech'
Note
digit_policy folds digit variants before the key is built, and the builder's
own confusable fold runs under the same policy (#896). "numeric" is the default
and a genuine no-op; "tr39" reaches more spoofs but destroys the numeric reading
of Arabic, Persian, Indic and Thai digits; "preserve" keeps a non-Latin numeral
in its own script (#949). See canonicalize for the measurements and the trade
(#885).
The output can be the empty string (#728).
Measured at Unicode 15.0.0, 140,200 single
characters reduce to "" here (2,732 excluding the Private Use
Area), and so does every string built from them. A caller keying a table
on this has all of them, plus "no value", competing for one slot.
There is no on_empty here: this returns text rather than a key. The
four key builders take one.
Pipeline steps¶
resolve_deletions → policy_pre_fold → NFKC → strip_zalgo(0) → strip_bidi → strip_zero_width → strip invisibles (#413) → confusables → strip_accents → strip_control → collapse_whitespace → NFC
from disarm import strip_obfuscation
# Homoglyphs (Greek/Cyrillic) folded, bidi override removed, emoji expanded.
# The emoji is left where it stands, not named (#910): a comparison surface must not
# insert attacker-chosen words. Use `demojize()` when the name is what you want.
assert strip_obfuscation("Ηеllо\u202eWоrld \U0001f600") == "HelloWorld \U0001f600"
# Strips ALL combining marks (zalgo and accents) but preserves case.
assert strip_obfuscation("Cáfé") == "Cafe"
Maximum-strength deobfuscation for content moderation, anti-phishing, and spam/NLP preprocessing. Strips every combining mark (zalgo and accents), resolves homoglyphs by TR39 visual similarity (Cyrillic р→p, not phonetic р→r). Leaves emoji where they stand rather than naming them (#910): a comparison surface must not insert attacker-chosen words into the value being compared — use demojize() when the name is what you want. Preserves case — case is meaningful, not deception. Does not transliterate; chain transliterate() on the result if you also need phonetic romanization.
PRESETS¶
from disarm import PRESETS
Dict mapping preset function names to their ordered pipeline steps. Each value is a list of (step_name, parameter) tuples in execution order.
assert PRESETS["canonicalize"] == [
("resolve_deletions", None),
("policy_pre_fold", "latin"),
("normalize", "NFKC"),
("strip_bidi", None),
("strip_invisibles", "comparison"),
("strip_control", None),
("strip_zero_width", None),
("collapse_whitespace", None),
("drop_repeated_marks", None),
("strip_zalgo", None),
("normalize", "NFC"),
("fixed_point", "confusables(latin) -> normalize(NFC)"),
("drop_repeated_marks", None),
("strip_zalgo", None),
]
assert PRESETS["canonicalize_strict"] == [
("resolve_deletions", None),
("policy_pre_fold", "latin"),
("normalize", "NFKC"),
("strip_bidi", None),
("strip_zero_width", None),
("strip_control", None),
("strip_invisibles", "comparison"),
(
"fixed_point",
"fixed_point(confusables(latin) -> normalize(NFC)) -> strip_cross_script_marks",
),
("drop_repeated_marks", None),
("strip_zalgo", None),
("collapse_whitespace", None),
("normalize", "NFC"),
]
Use PRESETS to audit exactly which transforms a preset applies. It is a mirror of the
step lists in src/presets/text.rs and src/presets/keys.rs, and a test reads those lists and
fails when the two differ — it drifted for a long time before that, missing steps that ran
and listing one that did not (Finding 5 of the Lean model in formal/lean/Presets).
Most names are the TextPipeline step or public function of the same name. Five are not:
| step | what it does |
|---|---|
policy_pre_fold |
Nothing under the default digit_policy. Under "tr39" or "preserve", the whole Latin confusable fold on the raw text (#885, #896) |
fixed_point |
Runs the inner steps, named in its parameter, as a group until the text stops changing (bounded) |
drop_repeated_marks |
Drops a nonspacing mark repeated on one base (UTS #39 §5.4, #835) |
strip_cross_script_marks |
Drops a combining mark whose script differs from its base's (#615) |
prototype_fold |
I to l, and under "tr39" 1 to l and 0 to O (#650) |
Two properties are not steps, so the lists cannot show them: search_key and sort_key
run their whole list to a fixed point under a non-default digit_policy, and every preset
raises ResourceLimitError when a step leaves the text more than 10 MiB longer than its
input (#768). A TextPipeline has neither fixed_point nor the preset-only steps, so it
reproduces a preset only approximately.
digit_policy on the key builders¶
On canonicalize, canonicalize_strict and strip_obfuscation, which fold confusables
anyway, a non-default policy changes the digit rows the fold reads. On catalog_key,
search_key and sort_key it does more: policy_pre_fold runs the whole confusable
table on the raw text, before transliteration, and search_key and sort_key have no
fold of their own at all under the default. So under "tr39" or "preserve" a Cyrillic
spelling of paypal keys as paypal rather than raural, and search_key and
sort_key rewrite |, " and ` as the other folding surfaces do.
from disarm import search_key
cyrillic_paypal = "".join(map(chr, (0x440, 0x430, 0x443, 0x440, 0x430, 0x6C)))
assert search_key(cyrillic_paypal) == "raural"
assert search_key(cyrillic_paypal, digit_policy="preserve") == "paypal"
assert search_key("a|b") == "a|b"
assert search_key("a|b", digit_policy="tr39") == "alb"
"preserve" is named for what it does to numerals, not for leaving the rest of the key
alone.
None here is a parameter, not an off switch
A None in the second position means the step takes no parameter, or runs at its
own default — the step is in the list, so it runs. This is the opposite of what
None means as a TextPipeline keyword, where it omits the step (#958). The
difference bites on strip_zalgo: ("strip_zalgo", None) above is a live step at
the default cap, while TextPipeline(strip_zalgo=None) compiles no step, and
TextPipeline(strip_zalgo=0) compiles a step that strips every diacritic. See
TextPipeline.
Policy Profiles¶
Named policy profiles provide pre-configured TextPipeline instances for common institutional and application workflows.
get_pipeline¶
from disarm import get_pipeline
pipe = get_pipeline("scholarly_cyrillic_iso9")
assert pipe("Москва") == "moskva"
Returns a fresh TextPipeline configured for the named profile. Raises DisarmError for unknown profiles.
A profile runs its steps again until the output stops changing (bounded), so calling it on
its own output returns that output. One pass was not always enough (Findings 3 and 4 of the
Lean model in formal/lean/Presets): the mark strip runs before the confusable fold and
before strip_pua, and the control and zero-width strips run after normalize, so
llm_guardrail kept a negation overlay on a symbol and then stripped it once the fold had
made the symbol a letter. A TextPipeline built from the same flags runs its steps once;
the two agree wherever one pass is already a fixed point, which includes every single code
point.
from disarm import get_pipeline
guardrail = get_pipeline("llm_guardrail")
cent_negated = chr(0xA2) + chr(0x338)
assert guardrail(cent_negated) == "c"
assert guardrail(guardrail(cent_negated)) == "c"
list_profiles¶
from disarm import list_profiles
print(list_profiles())
# ['code_context', 'library_catalog_key_eu', 'llm_guardrail', 'ml_corpus_normalize',
# 'normalize_web_input', 'rag_ingest', 'scholarly_cyrillic_iso9', 'search_index']
Returns sorted list of available profile names.
Available profiles¶
| Profile | Steps | Output |
|---|---|---|
code_context |
strip_bidi → strip_control → strip_zero_width | UTF-8 |
scholarly_cyrillic_iso9 |
NFKC → strip_plane14 → transliterate (ISO 9) → fold_case → strip_control → strip_zero_width → strip_pua → collapse_whitespace | UTF-8 |
library_catalog_key_eu |
NFKC → strip_plane14 → strip_accents → transliterate → confusables → fold_case → confusables → fold_case → strip_control → strip_zero_width → strip_pua → collapse_whitespace | ASCII |
normalize_web_input |
NFKC → confusables → strip_control → strip_zero_width → strip_pua → collapse_whitespace | UTF-8 |
ml_corpus_normalize |
NFKC → strip_plane14 → demojize → strip_accents → fold_case → strip_control → strip_zero_width → strip_pua → collapse_whitespace | UTF-8 (no transliteration: a script without accents keeps its letters) |
search_index |
NFKC → strip_plane14 → strip_accents → transliterate → fold_case → strip_control → strip_zero_width → strip_pua → collapse_whitespace | ASCII |
llm_guardrail |
resolve_deletions → NFKC → strip_zalgo(0) → strip_bidi → strip_plane14 → strip_accents → confusables → fold_case → confusables → fold_case → strip_control → strip_zero_width → strip_pua → collapse_whitespace | UTF-8 |
rag_ingest |
resolve_deletions → NFKC → strip_bidi → strip_plane14 → strip_accents → transliterate → strip_control → strip_zero_width → strip_pua → collapse_whitespace | ASCII |
Each list is what the profile's steps reports, and every profile runs it to a fixed point.
llm_guardrail hardens text against prompt-injection and homoglyph/zalgo/bidi obfuscation before it reaches an LLM (digits are never remapped to letters). rag_ingest canonicalizes documents for retrieval pipelines while preserving case.
Homoglyph handling: rag_ingest romanizes, it does not visually-fold (#258)
The two guardrail profiles canonicalize homoglyphs differently, and the distinction matters for spoof resistance:
llm_guardrailrunsconfusableswithouttransliterate, so a Cyrillic look-alike of "paypal" (раураl) is visually folded topaypal— it collides with the real Latin term (good for "treat the spoof as the word it imitates").rag_ingestrunstransliterate, which phonetically romanizes the same input toraural— a distinct key, so the spoof does not impersonate the real term, and legitimate non-Latin text still romanizes for retrieval (Москва → Moskva).
These are deliberate trade-offs of the fixed step order (transliterate runs
before confusables; running confusables first would mangle legitimate
Cyrillic/Greek into mixed-script gibberish). Adding confusables to
rag_ingest would be a no-op — transliterate has already consumed the
non-Latin characters. If you need homoglyph spoofs folded onto the term
they imitate, use llm_guardrail (or a dedicated confusables pass), not
rag_ingest.
See Policy Templates for detailed usage guidance and institutional recipes.