XBRL and Structured Filing Data
Filings are tagged so that individual figures can be extracted by machine. It makes comparison across thousands of companies possible, with a specific set of caveats.
MadStockAlerts Research · Updated August 28, 2026
What to take away
- Financial statement figures in filings are tagged with standard identifiers.
- The tagging makes bulk extraction and comparison possible.
- Companies may use custom tags where no standard one fits.
- Custom tags are where comparability breaks down.
- The SEC publishes bulk datasets built from the tagged data.
MAD Academy Training Video · 0:45
Filings a Machine Can Read
Every US filing is tagged in a structured format, which is why screening across thousands of companies is possible at all.
This lesson is part of a Stock Alerts + Tools plan.
What tagging does
Each figure in a financial statement carries a machine-readable tag identifying what it is: revenue, cost of revenue, cash and equivalents, and so on, drawn from a standard taxonomy. A program can then extract the same line from thousands of filings.
That is what makes any screening or comparison across the whole market possible at all. Without it, every comparison would require someone to read every filing.
Where it breaks down
- Custom tags, or extensions, are permitted where no standard element fits, and they are not comparable across companies.
- The same economic item can legitimately be tagged with different standard elements by different filers.
- Tagging errors occur, including sign errors and misplaced scale factors.
- Restatements and amendments produce multiple tagged versions of the same period.
The first item is the substantive limitation. A company with a large number of custom extensions is one whose figures cannot be compared against peers by machine, and the extension count is itself a measurable property of a filing.
- 1The company prepares the filingAnd tags each figure
- 2Standard elements where they fitComparable across filers
- 3Custom extensions where they do notNot comparable to anything
- 4Extraction into a datasetWhich inherits every choice above
What is published
| Resource | What it contains |
|---|---|
| Financial statement data sets | Quarterly bulk extracts of the tagged numeric data |
| Company facts | Every tagged fact for one company, across its filings |
| Frames | One concept across many companies for one period |
| The submissions record | A company's filing history in machine-readable form |
All of these are published free and without registration. The frames endpoint in particular answers the comparison question directly: one concept, one period, every company that reported it.
The caution that matters most
A number extracted correctly can still be the wrong number. Different companies legitimately classify the same item differently, and a machine comparison inherits every one of those choices without seeing them. The filing remains the authority, and the dataset is a way of deciding which filings to open.
What can be built from it
The practical consequence of tagged filings is that questions covering the whole market become answerable without reading anything, subject to the caveats above.
- Screening on any reported line item across every filer, rather than on a data vendor's derived fields.
- Tracking how a single company's disclosure of one item has changed across years, including restatements.
- Comparing an item's definition across an industry by examining which tags were used.
- Counting custom extensions, which is itself a rough measure of how comparable a filer's data is.
The second item is worth emphasising. Because each filing's tagged data is preserved, the figure as originally reported remains available alongside any later restatement, which is exactly the point-in-time property that backtesting requires.
That connects this article directly to the look-ahead bias article: the reason restated figures are a hazard in testing is that most data sources overwrite them, and the tagged filing record does not.