INFLXD MediaSubscribe →
Case Study

How AlphaSense turned a licensed research corpus into a generative answer engine

The Generative Search rollout became the reference case for layering LLMs over a permissioned content perimeter without breaking publisher or expert-network contracts.

INFLXD Research··7 min read
How AlphaSense turned a licensed research corpus into a generative answer engine

AlphaSense's June 2023 launch of Generative Search is the cleanest public example of a primary-research aggregator converting a licensed transcript and filings library into a citation-anchored LLM product. The rollout answered a specific question the buy-side was asking through 2023: whether a content platform holding third-party broker research, filings, and expert-call transcripts under license could layer generative AI over that corpus without breaching publisher terms or creating MNPI exposure. Two years and one Tegus acquisition later, the feature sits at the center of a company that has crossed $600M ARR and a $7.5B valuation.

Background: the 2023 buy-side stalemate

By early 2023, most large buy-side firms had years of accumulated primary research sitting inside AlphaSense: broker notes, 10-Ks and 10-Qs, earnings transcripts, news, and, at firms that also subscribed separately, expert-call transcripts from providers such as Tegus, Third Bridge, and GLG. The workflow to get value out of that corpus was still keyword search, filtered by ticker, date, and source type. Analysts opened documents one at a time and read for the answer.

In parallel, those same analysts were quietly opening ChatGPT in another tab. The appeal was obvious: paste a paragraph, ask for a summary, get a coherent answer in seconds. The problem was equally obvious. General-purpose LLMs had three disqualifying properties for a compliance-reviewed research workflow. They hallucinated numbers. They could not cite the source of a claim. And pasting a licensed broker note or an expert-call transcript into a public model was, at minimum, a violation of the publisher's terms of use, and at worst a data-leakage event that a compliance team would treat as a firing offence.

The research-technology industry's question through the first half of 2023 was whether any of the incumbent aggregators could close that gap. The vendor that could layer generative AI over a permissioned corpus, with citations back to the source document, and without breaking the contracts that let it host the content in the first place, would own the answer layer for buy-side research.

The approach: build generative on top of retrieval, not next to it

AlphaSense's answer was Generative Search, announced in June 2023 through the product blog. The design choices are worth reading closely because they are the choices the rest of the category has since converged on.

First, the feature was scoped to the existing content universe. AlphaSense described it as generative summarization across more than 10,000 sources already licensed into the platform: broker research through the Wall Street Insights partnership, SEC and global filings, company presentations, news wires, and trade press. The generative layer did not reach outside that perimeter. An analyst asking Generative Search a question got an answer drawn only from documents the firm's subscription entitled it to see. That single design decision resolved most of the publisher-contract question up front. The LLM was reading over licensed content the user was already permissioned to read.

A dense stack of stamped, permission-sealed transcript pages compressed under a glowing prompt-bar cursor, the pages fanning out at the bottom into neatly numbered citation chips that snap back to the

Second, the feature was built on top of the existing Smart Synonyms search infrastructure rather than as a separate product. Retrieval came first; generation was a summarization step over the retrieved set. Every generated sentence in the answer pane was linked back to the specific source document and page it came from, so an analyst could click through and verify the underlying text. That is the citation-back-to-source UX that has since become table stakes for enterprise research AI, but in June 2023 it was the differentiator.

Third, CEO Jack Kokko positioned the product publicly as ChatGPT for financial research with enterprise-grade permissioning and source attribution. The framing addressed the two objections that had blocked general-purpose LLMs from the buy-side seat: hallucination, answered by the citation UX, and provenance, answered by the licensed-corpus scope.

Extending the perimeter: the Tegus deal

Generative Search worked on the content AlphaSense already had. The strategic question the rollout raised almost immediately was what happened when the perimeter itself expanded.

The answer came in September 2024, when AlphaSense acquired Tegus for $930M. Tegus at the time held roughly 150,000 expert-call transcripts, built up over years of primary research calls conducted through its own moderator and compliance workflow. Folding that archive into the AlphaSense corpus put a proprietary expert-network transcript library, licensed broker research, filings, and news under a single LLM answer layer. As of the announcement, AlphaSense was the first content aggregator to hold all four categories in one permissioned index that a generative surface could read across.

The compliance shape of the combined product is worth noting. Expert-call transcripts are the highest-sensitivity content type in the primary-research stack. They are the layer closest to potential MNPI exposure and the layer where publisher and expert-network terms are most restrictive about downstream reuse. Combining that archive with a generative summarization layer only works if the underlying content was captured under terms that permit it, and if the output UX keeps every generated sentence traceable back to the original transcript passage a compliance reviewer can audit. AlphaSense's existing citation-back-to-source pattern from Generative Search extended to the Tegus content on the same terms.

Outcome: AI features as the commercial anchor

By mid-2025 AlphaSense was disclosing that AI features were driving roughly half of new-business conversations, according to reporting around its subsequent funding round. The company raised $650M in a Series F in May 2024 at a $4B valuation, then a further round in 2025 that took the valuation to $7.5B on more than $600M of ARR, per Reuters. The Generative Search feature became the anchor for a broader product line, including the Enterprise Intelligence and Assistant products that extended the same generative pattern to a firm's own internal documents alongside the licensed external corpus.

The commercial pattern that emerged is instructive. Analysts did not buy Generative Search as a standalone AI product. They bought AlphaSense, and Generative Search was the reason the platform sat above the alternative of running keyword search over the same content or pasting excerpts into a general-purpose LLM. The AI layer converted a content-aggregation subscription into a research-workflow subscription, which is a different, stickier product with a different price point.

What it signals for the industry

The Generative Search rollout is now the reference architecture for how a research aggregator or an expert network converts a licensed content library into a generative product. Three structural lessons stand out.

The moat is the content perimeter, not the model. Any competent engineering team can wire a modern LLM to a retrieval index. What most teams cannot do is assemble a decade of publisher, broker, and expert-network licensing agreements that permit that retrieval index to exist in the first place. AlphaSense's defensibility comes from the contracts, not the code. This is the same lesson Bloomberg terminal, FactSet, and S&P Capital IQ have taught for thirty years, applied to the generative era.

Citation UX is the compliance product. The reason Generative Search passed buy-side procurement in 2023 while general-purpose LLMs did not is that every generated claim linked to a specific source page a compliance reviewer could audit. Any research-AI product that ships without that traceability is not sellable into a compliance-reviewed workflow, regardless of how good the underlying model is.

Expert-call transcripts are the highest-value corpus in the stack. The Tegus acquisition placed a $930M valuation on roughly 150,000 primary-research transcripts. That is a public price signal about what buy-side firms are willing to pay to have a proprietary, permissioned, LLM-ready expert-network archive under the same answer layer as their filings and broker research. Expert networks holding transcript archives of comparable scale should read the Tegus deal as the market-clearing price for that asset class in an AI answer-layer world.

AI disclosure: This article was produced with AI assistance and may contain inaccuracies. If something looks wrong, email media@inflxd.com and we will review and correct it promptly. See our Terms.

Position B disclosure: INFLXD has commercial relationships with one or more of the companies named in this article. See our editorial disclosures.

From INFLXD

Powering institutional-grade transcription for expert networks.

INFLXD provides AI-powered, human-edited transcription with sub-1% error rates for the world's leading expert networks and financial research firms.

Visit inflxd.com →