INFLXD MediaSubscribe →
Analysis

The agent-readable disclosure layer: how 10-K filings became the template for every research artifact

The SEC's structured-data stack is quietly setting the ingestion contract for buy-side AI, and the rest of the research corpus is being pulled toward it.

INFLXD Research··12 min read
The agent-readable disclosure layer: how 10-K filings became the template for every research artifact

The SEC spent the better part of a decade mandating that public companies tag their financial disclosures in a machine-readable format almost no human ever reads. That work is now the upstream template for how buy-side AI agents consume the entire research corpus, and the pressure is traveling downstream fast.

Inline XBRL, the tagging standard the agency maintains through its structured data program, was built for a world of regulators, academic researchers, and a small cohort of quant funds. The large-language-model wave inherited it by accident. When a 10-K is already tagged at the line-item level, an agent does not need to parse a PDF table, infer a column header, or guess whether a number is in thousands or millions. It queries a schema. That difference, boring on its own, is reshaping the economics of every adjacent research artifact: earnings transcripts, expert-call libraries, broker research, management commentary.

The SEC built an ingestion layer before anyone asked for one

Inline XBRL is not new. The specification has existed in production form since the mid-2010s, and the SEC phased in the mandate across large accelerated filers, smaller reporting companies, and eventually funds and foreign private issuers over roughly a decade. The idea was straightforward: embed the structured tags inside the human-readable document rather than publishing two separate artifacts. A line item for revenue carries a tag that identifies it as us-gaap:Revenues, specifies the period, the currency, the decimals, and the context. The number a reader sees on the page and the number a machine extracts are, by construction, the same number.

For most of that history, the primary consumers were the SEC's own economic analysis division, a handful of academic groups, and the systematic quant desks that had the engineering resources to build and maintain their own XBRL parsing pipelines. Everyone else used Capital IQ, FactSet, or Bloomberg as the intermediation layer and treated the raw filings as a human-reading problem.

The LLM wave broke that intermediation layer open. A general-purpose agent cannot afford a per-ticker Capital IQ subscription for every query it runs, and the economics of re-OCR-ing a 300-page 10-K at inference time are brutal. The structured feed, which the SEC publishes free and in a format designed for programmatic retrieval, is the obvious substrate. The agent fetches the tagged document, resolves the schema, and answers the quantitative question against the source. The unstructured prose, management discussion, risk factors, notes to financial statements, becomes the context layer around a quantitative core that is already resolved.

The shift is subtle in any single query and compounding across a research workflow. An analyst who runs forty quantitative lookups a day no longer waits on a model to re-read a filing it has read a hundred times. The lookup resolves against a schema. The prose model is reserved for the questions that actually require reading.

EDGAR Next formalizes what the agents were already doing

EDGAR Next, the SEC's modernization of the filing system itself, matters less for what it changes about the filings and more for what it changes about access. The program introduces account-based authentication, API-first submission paths, and a general tightening of the technical contract between filers and the agency. For an AI vendor, the signal is that the SEC is treating its filing corpus as infrastructure, not as a document library.

That framing has downstream consequences. When the regulator publishes a stable, versioned, machine-readable API with predictable schemas, the cost of building a differentiated ingestion layer collapses. The moat moves one level up: from who can parse the filings to who can reason over them, who can cross-reference them against transcripts and expert calls, and who can present the output in a form an investment committee will accept.

A single 10-K page rendered as a precision-cut key, its teeth formed from XBRL bracket-tags, sliding into the ingestion port of a glowing terminal panel ,  while a stack of untagged research PDFs waits

We read EDGAR Next as the regulator doing the undifferentiated heavy lifting that every AI vendor would otherwise duplicate. That is good for the vendors in aggregate and bad for any vendor whose pitch rested on proprietary ingestion of public filings.

The taxonomy keeps extending

The tagged surface is not static. The FASB 2024 GAAP Taxonomy release added and refined elements across revenue, income taxes, and crypto-asset disclosures, among other areas. The SEC's rulemaking has continued to pull additional disclosures into structured form: 10-K cover pages carry tagged elements, share repurchase disclosures under 2023 rulemaking are tagged, and the climate-disclosure rule finalized in 2024, currently stayed pending litigation, was drafted with structured tags built in.

The directional trend is clear even if any single rule's fate is not: new categories of disclosure arrive pre-tagged. A climate rule that eventually clears litigation, or a successor rule in a different administration, will carry the same contract. An agent that is already built against the taxonomy inherits the new disclosures for close to zero marginal engineering cost.

This is where the structured-disclosure stack starts to look less like a filing format and more like a slow-moving standard for how companies describe themselves to machines. The GAAP taxonomy covers the income statement, the balance sheet, the cash flow statement, and large portions of the notes. The SEC's extensions cover governance, repurchases, and (prospectively) climate. The remaining gaps are what the rest of the research corpus has to fill.

Where the tagged surface stops

What is not tagged, and will not be tagged in any reasonable near-term rulemaking:

  • Management's qualitative commentary during earnings calls
  • Analyst questions during the Q&A segment
  • Industry expert interviews conducted through expert networks
  • Broker research, sell-side notes, and desk commentary
  • Private company data, pre-IPO disclosures, and most of the alternative-data universe

The gap is the opportunity. An agent that has resolved the quantitative core of a 10-K through XBRL still needs to answer questions like what did management actually say about pricing, what did the industry expert say about competitive dynamics, and what are the sell-side desks hearing from channel checks. Those questions live in prose that was never tagged by any regulator.

Why the research corpus is being pulled toward the same contract

The buy-side AI vendors operating against SEC filings, Rogo, Hebbia, Bridgetown Research, Maywood, OpenAI's ChatGPT for Financial Services, among others, are not building bespoke parsers for every data source. They are building agent runtimes that expect data to arrive in a form they can query. When one corner of the corpus (the filings) arrives pre-tagged, pre-chunked, and pre-schematized, the pressure on every other corner to match that contract is immediate.

Consider the earnings call transcript. In its raw form it is a prose document with speaker labels, often inconsistently applied, no stable identifiers for the speakers across calls, no segment tags linking a management comment to the business segment it references, and no durable timestamps that an agent can use to resolve back to an audio source. Compared to a tagged 10-K, it is a mess.

For the agent, that difference shows up as cost. Resolving a question like how did management discuss segment X margin pressure in the Q3 2025 call against an untagged transcript requires the model to read the whole transcript, infer segment attribution from context, and attribute statements to speakers whose names it has to disambiguate from a header block. Against a transcript with speaker IDs, segment tags, and timestamp anchors, the same question resolves against a schema lookup followed by a short prose summary.

The transcript that matches the tagging contract wins agent-runtime placement; the transcript that does not gets re-processed by the agent itself, re-chunked, re-attributed, and treated as a lower-trust source for anything the agent cannot verify.

The same pressure applies to expert-call transcripts, where speaker attribution carries additional weight because the expert's identity, employment history, and compliance status are part of what makes the content usable. It applies to broker research, where segment tagging and ticker resolution are the difference between a note that an agent can cite and a note the agent treats as background prose.

Three paths for how this plays out

We read the structured-disclosure drift as having three plausible trajectories over the next two to three years. These are our scenarios, not forecasts.

The convergence case

Research vendors, including transcript providers, expert networks, and broker research platforms, converge on a shared tagging contract that mirrors the SEC's structured stack. Speaker IDs become durable across calls. Segment tags link commentary to GAAP taxonomy elements. Timestamps and provenance signatures travel with every chunk. The agent runtime treats all research artifacts as queryable schemas, with the prose as context. This is the scenario in which the research corpus becomes a genuine agent-first substrate, and the vendors that moved first win the runtime placement.

The fragmentation case

Vendors each build their own tagging layer, incompatible with each other and with the SEC taxonomy. Agents build adapter layers for the three or four largest vendors and treat the rest as prose. A handful of integration standards emerge, probably from the model providers themselves rather than from the research vendors, and the vendors that align with the dominant model provider's conventions capture the runtime. This is the messier outcome and arguably the more realistic near-term one.

The intermediation case

A new category of middleware emerges that sits between the research vendors and the agents, performing the tagging, chunking, and provenance work on the vendors' behalf. The vendors keep their prose-first production workflows; the middleware produces the agent-readable surface. This is close to what happened in the previous generation with Capital IQ and FactSet, except the consumer is now an agent runtime rather than a human analyst, and the margin structure is very different.

Our base case is a mix of the second and third: fragmentation at the vendor layer, with middleware filling the gap for the vendors that cannot or will not build the ingestion contract themselves. The convergence case requires a coordination outcome that the research industry has not historically produced.

Who's affected, and how the second-order effects travel

The pressure does not stop at the vendors directly selling to buy-side AI tools. It travels through the supply chain in several directions.

Expert networks are the clearest case. The content they produce, hour-long interviews with industry operators, is prose by nature and prose by production workflow. An expert call becomes a transcript, the transcript sits in a library, an analyst reads it or searches it. The agent-runtime version of that workflow requires the transcript to carry speaker identity, employment history at the time of the call, segment and ticker attribution for the companies discussed, timestamp anchors for provenance, and compliance signatures that let the agent verify what it is allowed to use. The networks that produce that surface get queried by agents. The networks that produce raw transcripts get re-processed by agents, which both costs the agent runtime and makes the content effectively lower-trust.

Earnings-transcript providers face a similar structural question. The leaders in this category have invested in speaker attribution and search for years, but the agent-runtime bar is higher: durable speaker IDs across calls, segment tags that link to GAAP taxonomy elements, provenance that lets an agent cite a specific utterance back to a specific second of audio. The transcript as a prose document is a solved problem; the transcript as a queryable schema is not.

Broker research has the hardest adaptation. The production workflow for a sell-side note is still fundamentally a human writer producing a PDF. The tagged surface, where it exists, is bolted on after the fact by distribution platforms. The agent-runtime version requires the research itself to be produced in a form the agent can query, which is a workflow change rather than a packaging change.

Private-company data vendors sit outside the SEC's regulatory perimeter but inside the same agent-runtime pressure. There is no XBRL for a Series C deck, but the agents still want to query private-company financials in the same shape they query public ones. The vendors that build taxonomies compatible with the GAAP structure, even informally, inherit the compatibility.

Compliance teams at the buy-side firms running these agents inherit a new problem: provenance for every claim an agent surfaces. When the quantitative claim resolved against an XBRL tag in a filed 10-K, the provenance is airtight. When the qualitative claim resolved against an expert-call transcript, the provenance is only as good as the transcript's attribution and compliance layer. The gap in trust between the two sources becomes a material issue for an investment committee.

What a research analyst should ask next

If we were briefing a research analyst preparing to interview an expert on this shift, these are the questions worth putting on the list:

  • Which of your firm's quantitative lookups now resolve against structured EDGAR data rather than against a vendor's parsed version, and what did that change in your query latency and cost per lookup?
  • How does your compliance team treat an agent-surfaced claim where the provenance resolves to an XBRL tag versus one where it resolves to an untagged transcript? Is there a different review bar?
  • Which of your research vendors have published a tagging and chunking contract for their agent-facing feeds, and which are still producing prose-only outputs?
  • For expert-call content specifically, what is the trust gap between a transcript with speaker attribution and segment tagging versus a raw transcript, and does that gap show up in which calls your analysts actually cite in a memo?
  • How are you thinking about private-company data in a world where the public-company corpus is pre-tagged? Are you building internal taxonomies, waiting for vendors, or treating private-company queries as a different workflow entirely?
From INFLXD

Powering institutional-grade transcription for expert networks.

INFLXD provides AI-powered, human-edited transcription with sub-1% error rates for the world's leading expert networks and financial research firms.

Visit inflxd.com →