The full corpus, as a data feed
Year-over-year language-change scores for 4,930 US filers —
deterministic, corpus-normalized, auditable to the underlying filings. Built for credit
monitoring, distressed and short research, and pre-deal diligence.
What you get
Per-company year-over-year 10-K scores (level + history), the live control ceiling,
currently-flagged status (476 companies today), and filing metadata.
10-Q coverage is in progress.
Methodology discipline
Scores are computed from filings as published (point-in-time; no restatement lookahead),
normalized against the whole corpus, and reproducible — every number maps to a real
function in the scoring code.
Methodology.
Delivery
REST API (JSON), CSV export, or a custom feed (S3/SFTP drop, bespoke schema, SLA) for
institutional use.
API docs.
What the feed is — and isn't: a transparent measurement layer, not a trading signal.
Every score is deterministic and reproducible from public EDGAR filings, and maps to a real function in the
scoring code. We do
not ship it as an alpha factor — we tested the language-stability quintile sort
ourselves, and while the in-sample ranking is real, it does not survive as a tradeable, all-cap return once
delisted filers are added back. Use the feed to
screen and prioritize filings for your own credit,
distress, or diligence work.
Methodology & limitations.
Research data — not investment advice.
Two ways in
The Desk plan ($499/mo) includes bulk API access, full-corpus
sweep alerts, unlimited watchlist, and 5 seats. For a custom feed, schema, or license terms,
talk to us directly.
Want a taste first? Free sample CSV ·
live flagged list ·
worked examples