API reference¶
The supported facade modules are listed below, with their public classes and functions generated from the source. These facades are the intended entry points for new code. Lower-level modules may remain importable, but a lower-level path does not become supported merely because Python can import it. The release guard validates the reviewed current facade. For guides, tutorials, and the per-script handbooks, see the project wiki; for what pyaegean is and where to begin, see the home page.
Where to start¶
aegean: the top-level namespace:load(),read_corpus(),combine(), the core value types, and the subpackages.aegean.core: the script-agnostic model (Corpus,Document,Token,TokenFormState,FormSegment,SourceMarkupRef,Sign, …); build your own withCorpus.from_records, slice withsubset, merge withmerge.aegean.greek: the Greek NLP pipeline. It covers:- text processing: normalize, tokenize, named sentence policies, scan, tag, lemmatize, and parse;
segment_text()/segment_sentences()with validated plugin results and exact boundary spans, andpipeline_tokens()for typed editorial forms;- isolated
GreekPipelineinstances with serializable configuration, boundediter_analyze_sentences()neural sentence streams, and explicit long-input and runtime-variant selection; - optional typed
TokenConfidence/SentenceConfidenceevidence throughconfidence_domainandconfidence_policy, and immutable annotation and domain registries throughlist_annotation_profiles()/list_domain_profiles(); - exact schema-4
AnalysisReceiptprovenance (runtime registry and evidence, composed output, post-processing, and optional calibration and policy hashes), manifest-validated neural bundles, and lossless CoNLL-U document I/O; - work discovery:
catalog()(the full 1,778-work index),popular_works(), andnt_books().
aegean.analysis: accounting reconciliation, sign-pattern search, statistics, comparison.aegean.scripts: the built-in writing-system plugins and their public facades for Linear A, Linear B, Cypriot, Cypro-Minoan, and alphabetic Greek.aegean.io: import your own text or token-carrier EpiDoc; export to EpiDoc, CSV, Parquet, RDF Turtle/JSON-LD, review tables, and the intentionally lossy Linear A Research Workbench format. It also exposes complete CoNLL-U envelopes, loss-aware spaCy/Stanza/CLTK adapters, and portable SHA-256-bound interoperability bundles. JSON and SQLite remain the full-fidelity corpus archives.
aegean.io interoperability facade¶
The reviewed public entry points are the typed core envelope (InteropDocument,
InteropTokenMetadata, InteropSentenceMetadata, InteropResult, and
InteropReport), CoNLL-U conversion (from_conllu, to_conllu,
from_token_records, from_ud_document), optional framework adapters
(to_spacy/from_spacy, to_stanza/from_stanza, to_cltk/from_cltk), and
portable InteropBundle helpers (bundle_from_document, dumps_interop_bundle,
loads_interop_bundle, read_interop_bundle, write_interop_bundle), and the
explicit CLTK seam (make_cltk_process). bundle_from_result, the sidecar
codecs, schema constants, and typed interoperability errors are also public and
are listed in the generated aegean.io reference. The adapter
dependencies are lazy and stay outside the core import path.
- aegean.db: SQLite round-trip persistence for a Corpus (stdlib-only, queryable rows + FTS5 search).
- aegean.mcp_server: the aegean-mcp Model Context Protocol server (the [mcp] extra).
API policy¶
The facade modules listed here are the supported entry points for new code. The
reviewed list of facade modules and explicitly selected symbols is kept in
scripts/api-manifest.json. Importability alone does not add a new implementation
module to the supported API. python scripts/check_api.py statically verifies
that the reviewed modules and symbols resolve. During pre-1.0 development through
the v4 Greek NLP segment, the facade may change directly when the design improves;
the CHANGELOG, manifest, docs, and tests must describe the resulting current API.
Build a corpus from your own text¶
aegean.io also reads: turn a string, a .txt file, a folder of texts, or a CSV into a
real Corpus with the full filter/query/analyse/export API. Greek text is run through the
Greek tokenizer; other scripts split on whitespace. Everything here is offline and
stdlib-only.
from aegean import io
corpus = io.from_text("μῆνιν ἄειδε θεὰ Πηληϊάδεω Ἀχιλῆος", doc_id="iliad")
print(len(corpus), "document(s),", sum(len(d.words) for d in corpus), "words")
# 1 document(s), 5 words
From the CLI, aegean import writes a corpus you can then analyse like any other:
$ aegean import myplato.txt -o myplato.json # --split whole|paragraph|line
wrote 1 document(s) to myplato.json
$ aegean stats myplato.json --top 5 # …then any corpus command works
Find a work to load¶
greek.catalog() is a bundled, offline index of every work with a Greek (-grc)
edition in Perseus canonical-greekLit + First1KGreek: 1,778 works, far beyond the 25
curated popular_works(). Each entry's id loads directly with greek.load_work
(metadata only: the texts stay fetched-on-demand, never bundled).
from aegean import greek
for w in greek.catalog(author="plato", source="perseus")[:2]:
print(w["id"], "—", w["title"])
# tlg0059.tlg001 — Euthyphro
# tlg0059.tlg002 — Apology
$ aegean greek catalog --author plato --source perseus -n 2
Greek works (36 matches)
┌────────────────┬────────┬───────────┬────────────────────┬─────────┐
│ id │ author │ title │ greek │ src │
├────────────────┼────────┼───────────┼────────────────────┼─────────┤
│ tlg0059.tlg001 │ Plato │ Euthyphro │ Εὐθύφρων │ perseus │
│ tlg0059.tlg002 │ Plato │ Apology │ Ἀπολογία Σωκράτους │ perseus │
└────────────────┴────────┴───────────┴────────────────────┴─────────┘
… and 34 more — narrow with --author/--title, or --limit 0 to list all (-o to save).
pip install pyaegean # core + Linear A + Greek (zero heavy dependencies)
pip install "pyaegean[all]" # bundled runtime extras, including neural (not Parquet/framework adapters)
See the README for the full extras matrix and Benchmarks for the Greek NLP accuracy numbers and protocol.