This library supports bidirectional conversion between Greek, Beta Code, and scientific transliteration. It provides predictable output, named standards-oriented presets, and diagnostics when a conversion loses information. Unknown characters are preserved instead of being silently discarded.
- Installation
- Quick start
- Choose the right API
- Use a preset
- Common recipes
- Detect information loss
- Extend or restrict the character inventory
- Guarantees and scope
- Advanced API
- Documentation
- Development
- License
The 1.0.0-beta.8 release is ESM-only and targets JSR and npm:
deno add jsr:@humanities/greek-conversion@1.0.0-beta.8
npm install @humanities/greek-conversion@betaImport the public module from JSR:
import {
betaCodeToGreek,
convert,
convertDetailed,
} from "@humanities/greek-conversion";The same package name and named exports are available from npm.
Use convert() for any pair of supported formats:
convert("a)/nqrwpos", "beta-code", "greek");
// ἄνθρωπος
convert("ἄνθρωπος", "greek", "transliteration");
// ánthrōpos
convert("ánthrōpos", "transliteration", "beta-code");
// a)/nqrwposDirectional helpers are also available when they make application code easier to read:
betaCodeToGreek("a)/nqrwpos"); // ἄνθρωποςThe supported format names are "greek", "beta-code", and
"transliteration".
Try the interactive playground
to compare formats, presets, information-loss diagnostics, and character
repertoires in the browser. It uses the library code from the current main
branch, which can be newer than the latest published package.
ASCII letter case has no semantic value in Beta Code input. Only * marks a
Greek capital, so A)/NQRWPOS and a)/nqrwpos both convert to ἄνθρωπος,
whereas *)/ANQRWPOS converts to Ἄνθρωπος. Marking every letter, as in
*)/A*N*Q*R*W*P*O*S, produces ἌΝΘΡΩΠΟΣ.
Canonical output follows the TLG placement and order rules:
- lowercase: letter, breathing, accent, iota subscript —
w(=|; - uppercase: asterisk, breathing, accent, letter, iota subscript —
*(=w|.
orthography.betaCodeCase selects "lowercase" or "uppercase" ASCII output
without changing the represented Greek letter case. The default and Perseus
preset use lowercase ASCII; tlg-core uses uppercase ASCII.
The engine-default yot spelling is j. Set orthography.yotBetaCode to
"#401", or use tlg-core, for the TLG character code.
| Need | API |
|---|---|
| Convert one string | convert() or a directional helper |
| Canonicalize or transform a string without changing its format | reencode() |
| Display a warning when information is lost | convertDetailed() |
| Reuse or inspect one parsed source | parse() and encode() |
All of these APIs use the same parser, canonical Greek document, conversion options, and encoders.
For a same-format conversion, reencode() avoids repeating the format:
reencode("λόγος", "greek", {
orthography: { finalSigma: "medial" },
});
// λόγοσPresets collect coherent options for a published convention or a practical interchange profile:
convert("Βίος Μπάλα", "greek", "transliteration", {
preset: "ala-lc-modern",
});
// Vios BalaAvailable presets:
| Preset | Target format | Intended use |
|---|---|---|
ala-lc-ancient |
transliteration |
Library romanization of Ancient and pre-1454 Medieval Greek |
ala-lc-modern |
transliteration |
Library romanization of post-1453 Modern Greek, including supported contextual digraphs |
bnf-core |
transliteration |
BnF/ISO-based romanization core for Ancient Greek cataloguing |
iso-843-type-1 |
transliteration |
ISO 843 Type 1 character transliteration for Ancient and Modern Greek |
perseus |
beta-code |
Lowercase-ASCII subset for Perseus and Morpheus interchange |
sbl-academic |
transliteration |
Scholarly Biblical-studies output retaining scientific diacritics |
sbl-general |
transliteration |
Reader-facing Biblical-studies output omitting most scholarly diacritics |
tlg-core |
beta-code |
Uppercase-ASCII TLG core for polytonic Greek interchange |
Custom options override only the corresponding fields of a preset:
convert("Βίος Μπάλα", "greek", "transliteration", {
preset: "ala-lc-modern",
orthography: { beta: "b" },
});
// Bios BalaThe priority is: library defaults, then preset options, then custom options. See the complete preset configuration table, including the precise scope and known limitations of each preset.
For repeated conversion with scope enforcement, bind the preset to a converter:
const perseus = createConverter({ preset: "perseus" });
const result = perseus.convertDetailed("ἄνθρωπος ϝ", "greek", "beta-code");
result.output; // a)/nqrwpos ϝ
result.diagnostics[0].code; // out-of-scope-characterOut-of-scope characters are preserved rather than silently converted. The dedicated character-extension documentation explains contextual ALA-LC numerals, custom extensions, and further restrictions.
For strict conversion, set outOfScopeBehavior: "reject" when creating the
converter. Both conversion methods then throw CharacterScopeError, whose
diagnostics identify recognized characters outside the repertoire or with
known undefined or unresolved preset mappings. For ISO Type 1 transliteration,
these include stigma, koppa, sampi, and the ambiguously identified archaic
koppa; see the preset repertoire audit.
This rejects recognized out-of-scope characters and known missing mappings;
it does not guarantee a reversible result. Inspect convertDetailed().losses
separately for distinctions lost by an in-scope conversion. The
preset behavior audit
records the remaining contextual limits.
Every effective default is available as an immutable, IDE-friendly object:
import {
DEFAULT_CONVERSION_OPTIONS,
resolveConversionOptions,
} from "@humanities/greek-conversion";
DEFAULT_CONVERSION_OPTIONS.orthography.longVowels; // "macron"
DEFAULT_CONVERSION_OPTIONS.unicode.composition; // "composed"
resolveConversionOptions({ preset: "sbl-general" });
// A complete ResolvedConversionOptions objectConversionOptions remains the partial input type. ResolvedConversionOptions
contains every effective field after applying defaults, a preset, and custom
overrides.
convert("Ἄνθρωπὸς ᾆ", "greek", "greek", {
orthography: { accentuation: "monotonic" },
});
// Άνθρωπός άThe transformation is mechanical: it does not use a lexicon or infer missing breathings. Polytonic Greek remains the default.
Use the shortcut to remove marks. In transliteration, it retains the h of a
rough breathing (ὁδός → hodos); in Greek and Beta Code, it removes the
breathing. To suppress the transliterated h explicitly, set
diacritics: { roughBreathing: "remove" }.
convert("ἄνθρωπος ᾆ ῑ", "greek", "transliteration", {
removeDiacritics: true,
});
// anthrōpos a iOr select semantic classes independently:
convert("ἄνθρωπος ἅγιος κἀγώ", "greek", "transliteration", {
diacritics: {
accents: "remove",
smoothBreathing: "remove",
roughBreathing: "preserve",
coronis: "remove",
},
});
// anthrōpos hagios kagōremoveDiacritics() exposes mark removal as a standalone helper.
Structural distinctions such as η → ē and ω → ō are retained.
With removeDiacritics: true, Greek source sigma forms (σ, ς, ϲ) are
retained by default, including when using convert(). Explicit sigma and
finalSigma options can still select another spelling.
Transliteration uses ASCII - by default; set
orthography: { hyphen: "typographic" } to output ‐ instead.
Use foldGreekVariants() to obtain a uniform Greek spelling for comparison:
import { foldGreekVariants } from "@humanities/greek-conversion";
foldGreekVariants("ϐίος λόγος ϲῶμα");
// βίοσ λόγοσ σῶμαThe helper maps medial beta ϐ to β, lunate sigma ϲ to σ, and final
sigma ς to medial σ. It parses and re-encodes Greek rather than relying on
NFKC, so the same semantic rules and numeral exceptions apply as during a
conversion. Diacritics and case are preserved unless their independent options
are supplied:
foldGreekVariants("ϐΊΟΣ", {
orthography: { letterCase: "lowercase" },
removeDiacritics: true,
});
// βιοσconvert(" ΦΙΛΗΒΟΣ\nΗΔΟΝΗ ", "greek", "transliteration", {
orthography: {
letterCase: "lowercase",
whitespace: "collapse",
},
});
// philēbos ēdonēCase values are "preserve", "lowercase", "uppercase", and "title".
Whitespace can be "preserve" or "collapse"; the latter trims the output
and replaces each Unicode whitespace run with one ASCII space.
convert("βηξφχυ", "greek", "transliteration", {
orthography: {
beta: "v",
eta: "ī",
xi: "ks",
phi: "f",
chi: "kh",
upsilon: "y",
},
});
// vīksfkhyOther policies cover long-vowel spelling, nasal gamma, modern digraphs,
systematic rh, double rho, medial beta, sigma style, contextual or uniform
final sigma, coronis, and alphabetic numerals. The selected spellings are
recognized on transliteration input when the same options are supplied.
Lunate sigma has two independent controls. sigma: "standard" | "lunate" | "preserve" selects its Greek and Beta Code glyph; "preserve" is useful for
mixed texts because the parsers remember whether each source sigma was lunate.
lunateSigma: "s" | "c" selects the transliteration of only those provenanced
lunate sigma graphemes:
const options = {
orthography: { sigma: "preserve", lunateSigma: "c" },
} as const;
convert("σϲς", "greek", "transliteration", options); // "scs"
convert("scs", "transliteration", "greek", options); // "σϲς"A global sigma: "lunate" policy remains stylistic: it does not make every
transliterated s become c. The bnf-core preset does not select this option.
longVowels accepts "macron" (the default) or "circumflex". It controls
only the structural representation of inherently long eta and omega. An
explicit macron alongside a circumflex, such as in ê̄, is treated separately
as a philological mark and is reproduced without changing the identified
letter.
archaicKoppa selects "k-dot-below" (the default, rendered ḳ) or "q"
for transliteration. The bnf-core preset uses "q" for both koppa forms, so
that distinction is intentionally lost on reverse conversion.
import {
formatGreekUnicode,
toUnicodeCodePoints,
} from "@humanities/greek-conversion";
formatGreekUnicode("ά;·", {
acute: "oxia",
questionMark: "greek",
anoTeleia: "greek",
});
// ά;·
toUnicodeCodePoints("ά;😀");
// ["U+1F71", "U+037E", "U+1F600"]Greek output supports composed or decomposed text, tonos or oxia, and explicit Greek punctuation scalars. These are representation choices, not linguistic transformations. See Greek accentuation and Unicode output.
Some conversions necessarily merge distinctions. convertDetailed() returns
the output and diagnostics from the same conversion request:
const result = convertDetailed("ἄνθρωπος", "greek", "greek", {
orthography: { accentuation: "monotonic" },
});
result.output; // άνθρωπος
result.lossy; // true
result.losses;
// [{ code: "removed-diacritic", ... }]Loss is evaluated on the canonical document. Unicode composition and
tonos/oxia are therefore not reported as destructive. Lunate-sigma provenance
is the exception among glyph distinctions: removing a known lunate form
reports removed-glyph-variant. Examples of other actual loss include removing
diacritics, decimalizing alphabetic numerals, and using context-dependent
spellings that merge distinct source sequences.
See the conversion-analysis contract and the information-loss matrix.
Create an isolated Converter when an application needs additional spellings,
opaque custom characters, or a restricted repertoire:
import { createConverter } from "@humanities/greek-conversion";
const converter = createConverter({
aliases: [{
letter: "theta",
format: "transliteration",
spellings: ["þ"],
}],
exclude: ["stigma", "koppa", "sampi"],
});
converter.convert("þeos", "transliteration", "greek"); // θεοςAliases inherit every rule of their built-in letter. New character definitions
only guarantee direct format mapping and case; they do not silently join Greek
diphthong, contraction, numeral, or diacritic rules. Excluded characters become
source-format literals, and converter.convertDetailed()
reports them separately from information loss. See character extensions and
repertoires.
- Every format pair is supported in both directions.
- For fixed formats and options, canonical output is stable and idempotent.
- Unknown literals are preserved.
- Greek punctuation and contextual Greek rules are handled semantically where the canonical document contains enough information.
- The engine does not infer a missing rough breathing and does not perform transformations that intrinsically require lexical or morphological knowledge.
- A reverse conversion is not necessarily lossless; use
convertDetailed()when that distinction matters.
The engine currently covers polytonic and monotonic accentuation, contextual diphthongs and breathings, crasis and coronis provenance, elision aspiration, sigma and beta variants, archaic letters, Greek punctuation, and marked alphabetic numerals. Presets are limited to the rules and characters implemented by the engine today.
Applications can work directly with the experimental canonical representation:
import {
encode,
parse,
validateDocument,
} from "@humanities/greek-conversion/document";
const document = parse("a)/nqrwpos", "beta-code");
const diagnostics = validateDocument(document);
const output = encode(document, "greek");Validation deliberately remains a separate diagnostic step instead of changing
the contract of convert(). It checks the interpreted document, not the syntax
of the original string; malformed source sequences can survive as literals.
The ./document entry point may still evolve during the 1.0.0 prerelease
series. See the document API contract
for ownership, transformations, and options, and validation
for the validation boundary.
| Topic | Document |
|---|---|
| Presets and exact option values | Presets |
| Greek orthography and Unicode | Greek Unicode |
| Detailed conversion results | Conversion analysis |
| Loss by format pair | Information loss |
| Canonical document validation | Validation |
| Canonical document API | Document API |
| Character extensions and repertoires | Character extensions |
| Preset character repertoire audit | Preset repertoire audit |
| Preset behavior audit | Preset behavior audit |
| BnF and ISO conformance corpus | Preset conformance corpus |
| ALA-LC conformance corpus | ALA-LC conformance corpus |
deno task check
deno task test
deno task site:buildThe last command builds the static playground in _site/. Serve that directory
with a local HTTP server to test the page; the Playground workflow publishes
the same build to GitHub Pages when Pages is configured to use GitHub Actions.
Copyright (C) 2021-2026 Antoine Boquet
Starting with 1.0.0-beta.7, greek-conversion is licensed under the
MIT License.
Previously published versions retain their original AGPL-3.0-or-later license.