Language Profiles¶
Functions for querying and extending transliteration language profiles.
list_langs¶
list_langs ¶
list_langs() -> list[str]
Return available language codes for transliteration.
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> "de" in list_langs()
True
>>> "ja" in list_langs()
True
Example¶
from disarm import list_langs
langs = list_langs()
assert langs == [
"am",
"ar",
"as",
"ban",
"bax",
"bg",
"bn",
"bo",
"bug",
"ca",
"chr",
"cjm",
"cop",
"cs",
"cy",
"da",
"de",
"dv",
"el",
"es",
"et",
"fa",
"fi",
"fr",
"ga",
"gu",
"he",
"hi",
"hr",
"hu",
"hy",
"is",
"it",
"ja",
"ja-kunrei",
"jv",
"ka",
"khb",
"km",
"kn",
"ko",
"lis",
"lo",
"lt",
"lv",
"ml",
"mn",
"mni",
"mr",
"mt",
"my",
"ne",
"nl",
"no",
"nod",
"nqo",
"or",
"pa",
"pl",
"pt",
"ro",
"ru",
"sa",
"sat",
"si",
"sk",
"sl",
"sq",
"sr",
"su",
"sv",
"syr",
"ta",
"tdd",
"te",
"th",
"tl",
"tr",
"tzm",
"uk",
"vai",
"vi",
"zh",
]
Returns both built-in and user-registered language codes, sorted alphabetically.
Tip
Use lang="auto" to auto-detect the language from the dominant non-Latin script in the input, instead of specifying a code manually. See Language Support for details.
lang_info¶
lang_info ¶
lang_info(code: str) -> LangMeta
Return metadata for a language code.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> lang_info("de")["name"]
'German'
>>> lang_info("cop")["script"]
'Coptic'
Returns a LangMeta dict with name, script, region, and context keys for
a language code. The context field is one of "none", "partial", or "full",
indicating how much the language benefits from context=True in transliterate().
Raises KeyError if the code is not recognized.
Example¶
from disarm import lang_info
meta = lang_info("de")
assert meta["name"] == "German"
assert meta["script"] == "Latin"
assert meta["region"] == "European"
assert meta["context"] == "none"
assert lang_info("cop")["script"] == "Coptic"
script_info¶
script_info ¶
script_info(script: str | Script) -> ScriptMeta
Return metadata for a Unicode script.
| Parameters: |
|
|---|
| Returns: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> script_info("Coptic")["default_lang"]
'cop'
>>> script_info(Script.THAI)["name"]
'Thai'
Returns a ScriptMeta dict with name, default_lang, example, and
context_aware keys for a Unicode script. Accepts either a script name string or a
Script enum value. Raises KeyError if the script is not recognized.
Example¶
from disarm import Script, script_info
meta = script_info("Coptic")
assert meta["name"] == "Coptic"
assert meta["default_lang"] == "cop"
assert meta["context_aware"] is False
# A Script enum value works too
assert script_info(Script.THAI)["name"] == "Thai"
list_context_langs¶
list_context_langs ¶
list_context_langs() -> list[str]
Return language codes that support context-aware transliteration.
These languages benefit from context=True in transliterate.
Each entry has a context field in its lang_info metadata
indicating the level of support: "full" or "partial".
| Returns: |
|
|---|
Examples:
>>> "ar" in list_context_langs()
True
>>> "de" in list_context_langs()
False
Returns the sorted list of language codes whose lang_info() metadata has a
context field other than "none" — i.e. the languages that benefit from
context=True in transliterate().
Example¶
from disarm import list_context_langs
assert list_context_langs() == ["ar", "fa", "he"]
assert "ar" in list_context_langs()
assert "de" not in list_context_langs()
register_lang¶
register_lang ¶
register_lang(code: str, mappings: dict[str, str]) -> None
Register or override a transliteration mapping for a language code.
Warning
This mutates process-global state consulted by every
transliterate/slugify/catalog_key/… call in the interpreter.
Treat it as startup-only / single-writer configuration: do not call
it from request-handling or library code in a multi-tenant process, where
it would silently alter every other caller's output. Call
seal_registrations after startup to make further changes raise.
A call already running when this one returns is not guaranteed a single
table: a batch transliterate(list) releases the GIL and reads the
table per character, so its items, and even one long string, can mix
the old mappings with the new. Calls that start afterwards see only the
new ones.
Note
Mappings keyed on ASCII characters do not apply to pure-ASCII input.
The core takes a fast path that returns all-ASCII text unchanged before
consulting language tables (ASCII is the transliteration target, so it
is normally identity). Language profiles are meant for non-ASCII source
characters (e.g. ä→ae). To remap an ASCII character, use
register_replacements instead — its keys run as a pre-pass that
executes ahead of the ASCII fast path and therefore do apply.
| Parameters: |
|
|---|
| Raises: |
|
|---|
Examples:
>>> register_lang("xx", {"Ä": "Ae", "ä": "ae", "Ö": "Oe", "ö": "oe"})
>>> transliterate("Ärger", lang="xx")
'Aerger'
Example¶
from disarm import register_lang, transliterate
register_lang(
"eo",
{
"ĉ": "cx",
"ĝ": "gx",
"ĥ": "hx",
"ĵ": "jx",
"ŝ": "sx",
"ŭ": "ux",
},
)
assert transliterate("ĉapelo", lang="eo") == "cxapelo"
# Verify registration
from disarm import list_langs
assert "eo" in list_langs()
Warning
This is a global, process-wide operation. Registered profiles persist for the lifetime of the Python process and are visible to all threads.
register_replacements¶
register_replacements ¶
register_replacements(replacements: dict[str, str]) -> None
Register global pre-transliteration replacements.
New entries are merged into the existing table. Existing keys are
silently overwritten. Use clear_replacements to wipe the
table, or remove_replacement to remove a single key.
Replacements are applied to the input as a left-to-right pre-pass before the main transliteration tables, using longest-match-at-each-position semantics (the longest registered key matching at a position wins, and its output is not re-scanned, so replacements never cascade). Keys may be multi-character and may be ASCII.
Warning
Like register_lang, this mutates process-global state shared
by every caller. Treat it as startup-only / single-writer configuration
and call seal_registrations afterwards in multi-tenant processes. As
there, a batch call already in progress can apply the old table to some
items and the new one to others.
| Parameters: |
|
|---|
Examples:
>>> register_replacements({"™": "(tm)"})
>>> transliterate("hello™")
'hello(tm)'
>>> clear_replacements()
Example¶
from disarm import register_replacements, transliterate
register_replacements(
{
"©": "(c)",
"®": "(R)",
"™": "(TM)",
}
)
assert transliterate("Hello™ World©") == "Hello(TM) World(c)"
Replacements are applied as a pre-processing step before the character-by-character transliteration lookup. They are global and persist for the process lifetime.
remove_replacement¶
remove_replacement ¶
remove_replacement(key: str) -> bool
Remove a single global pre-transliteration replacement by key.
| Parameters: |
|
|---|
| Returns: |
|
|---|
Examples:
>>> register_replacements({"©": "(c)"})
>>> remove_replacement("©")
True
>>> remove_replacement("©")
False
Example¶
from disarm import register_replacements, remove_replacement, transliterate
register_replacements({"©": "(c)", "®": "(R)"})
assert transliterate("©®") == "(c)(R)"
assert remove_replacement("©") == True
assert remove_replacement("©") == False
assert transliterate("©®") == "(c)(R)"
clear_replacements¶
clear_replacements ¶
clear_replacements() -> None
Clear all global pre-transliteration replacements.
Examples:
>>> register_replacements({"©": "(c)", "®": "(r)"})
>>> clear_replacements()
Example¶
from disarm import register_replacements, clear_replacements, transliterate
register_replacements({"©": "(c)", "®": "(R)"})
assert transliterate("©®") == "(c)(R)"
clear_replacements()
assert transliterate("©®") == "(c)(R)"
Note
clear_replacements() removes all user-registered replacements. Built-in transliteration tables are not affected.