About FilingDrift

What this is

FilingDrift is a language-change scoring tool for SEC 10-K annual and 10-Q quarterly filings. It measures how much a company's filing language changes year over year — the directed increase in distress vocabulary — normalized against the whole corpus, and flags the filings whose wording moved far more than normal.

It is a first-pass screen, not a forecast. A high score tells you a filing's language moved in a statistically unusual way — for the company and across the whole corpus that year — and is worth opening and reading. It does not predict returns, and it is not a trading signal.

Corporate distress often has a pre-crisis signature in language. CFOs rarely say "we're in trouble" outright — they gradually introduce hedging language, new risk-factor categories, and liquidity disclosures that weren't there before. SVB's 2022 10-K scored 57.5 against a 95th-percentile control ceiling of 51.5; the FDIC arrived 14 days after filing. That's one clean example of what the screen surfaces — the scorecard below shows the hits and the misses.

What this is not

  • Not investment advice. We score language. We don't predict price movements, recommend positions, or guarantee outcomes.
  • Not a crystal ball. Most of the crisis companies in our hand-labeled set exceeded the control ceiling before their collapse — but not all. Party City, Revlon, and SI were missed (Party City's distress vocabulary is common enough across the corpus that the corpus-wide weighting discounts it). FilingDrift is a screening tool, not a prediction one.
  • Not a financial advisory service. Latent Systems SAS is a software company. We are not registered investment advisors, broker-dealers, or credit rating agencies.
  • Not affiliated with the SEC. We use publicly available data from EDGAR. We have no relationship with the SEC or any regulatory body.
  • Not proprietary or insider information. Every score is derived entirely from public SEC filings available on EDGAR. We have no access to non-public information about any company, and no score reflects anything that was not already in the public record at the time of filing.

How the score works

The score measures two things independently:

  • What's new or escalating. Phrases that appeared for the first time, or dramatically increased, weighted by how rarely other companies use them. A phrase that only SVB was saying in 2022, that it hadn't mentioned before, scores higher than boilerplate that every bank uses.
  • Where the language is heading. We measure how the current year's language has shifted from last year, in the sections most predictive of distress — risk factors, liquidity disclosures, MD&A. This catches the drift that keyword lists miss.

Both are calibrated against the healthy companies in our corpus — currently 4930+ tracked. The 95th percentile of their filing-pair scores is the control ceiling (51.5). Scores above it are flagged.

The algorithm is deterministic: no AI generation, no prompting, no summarization. The same filing always produces the same score.

See the full FAQ →

How it did on known crises — the honest scorecard, hits and misses

The table below shows how the score performed against a hand-labeled set of crisis companies in our corpus. Events include bankruptcies, bank failures, FDIC seizures, and Chapter 11 filings (some companies subsequently emerged).

Read this as illustration, not a benchmark: it's a small, hand-labeled, in-sample set of ~43 companies. Most of the well-known failures below crossed the ceiling before their event; several did not — Party City, Silvergate (SI), and Revlon were missed — and we list the misses on purpose. A larger, out-of-sample evaluation is on the roadmap.
Company Event Peak score Lead time Result
PRTY (Party City) Bankruptcy 2023 46.3 Missed
NKLA (Nikola) Bankruptcy 2023 85.3 3.7 years Detected
BBBY (Bed Bath) Bankruptcy 2023 138.5 2.0 years Detected
RITEAID Bankruptcy 2023 79.2 167 days Detected
SVB Financial Bank collapse 2023 57.5 14 days Detected
SI (Silvergate) Liquidation 2023 15.1 Missed
REVLON, CHKAQ Various <2 No data †

† REVLON and CHKAQ (Chesapeake) have a single, sparsely-parsed filing pair in our corpus — insufficient history to compute a meaningful change score. We count them as misses to avoid cherry-picking. The score requires at least two consecutive filings to measure change. (Party City, by contrast, has full history but its distress vocabulary is common enough across the corpus that the corpus-wide weighting discounts it — a genuine miss, not a data gap.)

False positives: 8 of 30 stable reference companies generated above-ceiling scores at some point — dominated by large financials (JPM, RTX). Some occurred during the COVID disruption years (2020–2021), when corpus-wide language shifts reduced the discriminating power of the period normalization. Others (e.g. RTX) followed a major corporate merger that produced large language changes for structural reasons.

📈
Corpus is actively expanding
Currently tracking 4930+ companies. We ingest new filings continuously as they appear on EDGAR. Coverage and statistical power improve with each new filing cycle. Control ceiling: 51.5.

Known limitations

  • M&A distortions. When a company acquires another and consolidates filings, the combined entity may show language change that reflects the target's pre-existing disclosure style, not a genuine deterioration.
  • Sector shocks. In years with corpus-wide stress (2008–2009, 2020), nearly all companies spike together. The per-period normalization is less informative when distress language is the baseline across the whole corpus.
  • Regulatory language. Banks under formal regulatory agreements (cease-and-desist, memoranda of understanding) are required to use specific disclosure language. This language reads as distressed but may not indicate approaching collapse.
  • Binomial false-flag rate. With 4930+ companies tracked over multiple years, some will exceed the ceiling by random variation. Companies with 10+ years of history have more opportunities to spike. We are working on potential solutions, e.g. adaptive thresholds.

Who we are

FilingDrift is built by Latent Systems, a small team of ML researchers based in Paris. We all have PhDs in machine learning. Our research focuses on training embedding models and studying the geometry of the spaces they produce: how meaning is encoded in high-dimensional representations, and what structural properties of those spaces can be exploited for detection, classification, and anomaly scoring.

FilingDrift grew out of that work. The question was whether financial distress leaves a detectable signature in how a company's filing language changes over time, and whether that signature appears before prices move. The core signal is a directed phrase-frequency change normalized across the whole corpus (with a secondary sentence-embedding component drawn directly from our research on representation geometry).

We are not a hedge fund, a financial services firm, or a consultancy. FilingDrift is a research product of an independent research company.

Questions, feedback, and enterprise inquiries: hello@filingdrift.com

Browse examples → Read the FAQ

This site uses a session cookie for authentication. We also use Plausible Analytics, a privacy-friendly, cookieless tool that collects no personal data and requires no consent under GDPR. See our Privacy Policy.