Agentic Search Requires a New Payment Model for Publishers
Parallel Web Systems founder and CEO Parag Agrawal argues that AI agents will make the web’s human-centered search and advertising model untenable: they will query far more often than people, but cannot generate the clicks and conversions publishers rely on. His proposed alternative combines agent-oriented retrieval—ranked by task outcomes rather than human clicks—with a Shapley-value-inspired payment system that would compensate content owners for the value their material adds to an agent’s work.

Publishers need a reason to let agents use the web
Parag Agrawal’s central concern is economic: the web’s dominant business models were built around human attention, and they become unstable when the visitor is an agent.
Advertising, in his account, has been an unusually effective form of differential pricing. Google Search, Twitter, and much of the open web could be free because a relatively small number of high-value commercial interactions subsidized a much larger volume of less valuable activity. Most queries may lose money, he says, while a smaller set produces enough revenue to support the system.
That arrangement depends on people arriving, seeing pages, and occasionally taking monetizable actions. A publisher with a subscription business can observe a familiar funnel: a thousand human visitors may produce 20 monthly subscribers. The equivalent loop is not yet clear when nameless agents access the same material. A publisher may not know whether an agent’s use will create a subscription, displace one, or simply extract information without producing a measurable commercial outcome.
Agrawal expects this uncertainty to make access more restrictive. Content owners may continue to pay for SEO to attract human readers while blocking agents, even when the agent is acting for a human. The problem, as he frames it, is not that the visit has no value. It is that the publisher cannot monetize the visit through the existing system.
The alternative available to the largest content owners has been a negotiated agreement with a model company, covering some combination of training data, inference-time access, and liability. Agrawal calls that a “head phenomenon”: most sites cannot strike bespoke deals with frontier labs. He also doubts fixed-fee arrangements will hold up even for the largest publishers. If inference grows sevenfold in one year and another sevenfold in the next, a two-year fixed-price deal does not scale with usage; the content owner’s share of the value declines by renewal.
Parallel’s proposed answer is a market in which content owners are paid for the value their information adds to agent work. The payment would vary both with the distinctiveness of the content and with the value of the work being done with it.
If you have unique differentiated data, you get paid more. If a banker in an expensive job reads your data versus my retired dad reads your data, the banker ends up paying more for that read.
This is the incentive problem Parallel is trying to solve alongside search itself: high-quality material needs a reason to remain available as agents become a meaningful share of web traffic.
Shapley values turn source contribution into a proposed payment
Agrawal’s framework for allocating payment is based on Shapley values. The question is straightforward: how much incremental value did a given source add to an agent’s output?
In principle, a system could remove one source from the available corpus, rerun the agent, and measure how much output quality declined. It could then test whether additional compute, a stronger model, or another source could recover that quality. If restoring the result without a particular source required an additional cent of compute, Agrawal’s intuition is that the source is worth roughly a cent.
Shapley values generalize that intuition to cases in which many contributors jointly create an outcome. If three parties collaborate to create something worth more than the sum of their individual contributions, the framework offers a way to divide that additional value so that each party has an incentive to participate. The literal calculation requires considering many alternate worlds: what happens when each possible combination of sources is present or absent.
That is not feasible to perform exactly across the web. Agrawal said a full Shapley calculation establishing that a content owner should receive a dollar could itself cost several dollars—more than the cost of running the agent and more than the payment being allocated.
Parallel’s intended approach is to approximate the calculation. It can run simulations in which URLs, domains, or collections of material are withheld; evaluate how the agent’s output changes; collect the resulting data; and train models to estimate contribution more cheaply. The theoretical foundation matters to Agrawal because he sees it as a test of incentive alignment: under conditions of perfect information, content owners and AI systems seeking information should both want to participate.
The proposed market would reward sources that are hard to replace and would price use differently depending on what the information enables. Agrawal compares that to the differential pricing that advertising supplied through human attention, but with the unit of value shifting from a person’s visit to a source’s contribution to an agent’s work.
He argues that the aggregate pool could be substantial. If organizations spend heavily on LLM inference for knowledge work, allocating roughly 2% to 10% of that spending to web data would, by his estimate, exceed today’s web-data business models outside major walled gardens such as Facebook and LinkedIn. He said Parallel has seen positive evidence from partnerships it has announced and estimated that, within 12 to 24 months, the approach could produce meaningful payments for a broad range of content owners.
Human clicks are the wrong feedback loop for agent work
Parallel’s technical premise follows from the economic one: agents are not merely another interface for conventional web search. They are a distinct customer with different behavior, different feedback signals, and potentially much higher search volume.
A conventional search engine crawls URLs, reads and organizes their contents into an index, then interprets a query and progressively narrows an enormous set of documents to a handful of results—or ideally one result. Agrawal describes that as a “billion-to-billion matching problem”: hundreds of billions of pages need to be matched to hundreds of billions of queries over time.
Historically, large search engines had two major advantages. They could afford full-web crawling and indexing, and they could observe vast quantities of human behavior. Clicks and ratings helped them learn whether one result was preferable to another.
Agrawal does not think click data transfers cleanly to agentic search.
Our view at Parallel is that human click data is a bug and agent doing work with search should rely on agent feedback, not human feedback.
A click captures part of a human browsing experience: whether a page appeared appealing, loaded quickly, and seemed useful enough for a person to open. An agent’s relevant outcome, in Agrawal’s framing, is whether retrieved information helped complete its task accurately, quickly, and within a cost budget. Parallel believes ratings data for that task can be produced more cheaply through expert evaluation and agent feedback than through the old system of accumulating human clicks.
The interface changes as well. Humans often type a few incomplete words, make typos, and rely on autocomplete to avoid specifying what they want. Agents can submit longer and cleaner requests with more context. That leaves less for the search engine to infer and allows it to solve a different retrieval problem.
The target is therefore not simply a more capable list of links. Parallel aims to return the highest-signal excerpts to an AI system’s context window, minimizing the token budget consumed by noise.
The agent can retrieve the filing instead of rewarding the summary page
Andrew Reed raised a practical implication of this distinction: many high-ranking human-search results are affiliate-heavy or SEO-optimized pages that appear to offer an answer but are not the underlying source.
Agrawal’s response was that those pages frequently solve a genuine problem for human users. Consider a request for a public company’s latest headline revenue number. The authoritative answer may appear in an SEC filing, perhaps on page 73 of a PDF. But the filing can be slow to load and tedious to navigate. A summary page may present the relevant figure immediately, above the fold, and load faster. It may be 99.99% accurate rather than the definitive source, but it has done useful work for a person who does not want to search through a long document.
That tradeoff between authority and convenience helped create a category of web pages optimized for human browsing. Agrawal argues that agents need not face the same constraint.
We are taking a excerpt from the most authoritative place on the web and trying to bring it to the agent’s context window.
The objective is to deliver a relevant passage from the authoritative source directly to the agent, rather than force the agent to click, load pages, and search through documents as a person would. The agent can then determine what to do next with the evidence.
Agrawal said that using Parallel Search will, in most cases, allow an agent to use under half the tokens it otherwise would, while becoming more accurate and faster end to end. If an agent is context- or memory-limited, using fewer tokens can make larger problems tractable. Otherwise, the same task can be completed more cheaply or quickly. In his account, agentic search is fundamentally an optimization problem across quality, cost, and latency: preserve useful signal while removing irrelevant material before it reaches the model.
Parallel sold research work before it had an instant index
A full web crawl and index is expensive, and customers often expect broad coverage from a search product immediately. Parallel’s answer was to avoid launching as a conventional search engine. It launched a search agent first.
The early product could crawl after a query arrived. For a deep-research task, a customer may tolerate a minute of waiting; Parallel has had products that spent as long as 10 minutes on research. That time can be used to crawl many pages, assuming the system has enough of a map to prioritize what it should retrieve.
An index, in Agrawal’s framing, is partly a latency optimization. A conventional engine crawls and preprocesses material in advance so that it can respond immediately when the query arrives. Parallel initially gave up that advantage. Rather than compete directly with Google’s response time on day one, its search agent competed with people already paid to gather and curate web information.
Early workloads included insurance underwriting and claims processing, sales-data enrichment, and finance tasks in which firms previously sent data-collection work overnight to human researchers before using it in a model. Parallel used those workloads to replace outsourced research and to gather empirical evaluations of what agents actually needed from search. As it served more customers, it could build an increasingly broad and sophisticated index.
Sonya Huang summarized the tradeoff as exchanging the crawl for inference-time compute. Agrawal agreed. The company initially prioritized quality and cost while treating latency as a dimension it could improve later.
Parallel’s output, as Agrawal describes it, is not a frontier model. He rejects the label “Neo lab” because the company is building a complement to models: infrastructure that makes a model-powered agent more useful. When another company ships a better model, he argues, Parallel should gain potential use cases rather than face a higher competitive hill.
The research problem is retrieval, ranking, compression, and compute allocation. A query is interpreted and enriched, then routed to different indexes: a broad index, a fresh index, and more structured collections that some people might describe as knowledge graphs. The system rewrites the query as needed for each index. Retrieval and successive ranking stages reduce tens or hundreds of billions of URLs and documents to smaller candidate sets, then to particular excerpts and paragraphs.
Different models, architectures, and features can be used at different stages. The endpoint is roughly a thousand tokens judged to contain the highest signal for the calling AI. Different versions of Parallel’s API allocate different amounts of compute along that path to meet different cost and latency constraints.
Agrawal said Parallel had previously allowed itself around three seconds for this process. Its Turbo product, he said, reduced that to 200 milliseconds. He characterized Turbo as the fastest, highest-quality agentic web-search offering on the market.
Parallel also announced an integration with Google Cloud as a search and grounding provider for enterprise agent APIs. Agrawal said developers building agents, chat applications, or other LLM inference products on Google Cloud can choose Google Search or product-integrated Parallel Search when attaching web search. He said Parallel worked with Google Cloud’s technical, training, product, and commercial teams to optimize use with Gemini models.
He does not take that arrangement as proof that model companies will necessarily own the whole search layer. Pretraining crawls and fresh web indexes serve different purposes, he argues. Model companies may not wait for slow-loading JavaScript pages or pursue difficult pages completionistically during training, because their token yield per unit of compute is poor. A fresh-search provider may see value in collecting precisely that material. At the same time, Agrawal considers agentic search a core adjacency to LLM inference: model companies can build the infrastructure, buy it, or partner for it.
Search volume rises when work becomes continuous
The thousandfold premise is not only about agents conducting deeper research than humans. It is also about who initiates the searches and how often.
A typical search agent may run five to 20 searches even when it produces an answer in seconds, Agrawal said. More capable configurations may execute hundreds or thousands. A single human request therefore becomes a multiplier on web activity.
The larger multiplier comes from background work. Agrawal offered the example of a lender with a portfolio of 10,000 small businesses. If people currently use web data once a month to assess changes in portfolio risk, an agent could run the process weekly, conduct many searches, produce a dashboard, and identify action items for human review. In that setting, a developer has turned one recurring workflow into hundreds of thousands or potentially a million searches.
Agrawal described a smaller version of the pattern in his own work. He uses a custom Notion agent that combines Parallel’s web APIs with internal company data to create meeting-preparation documents. After he created the initial prompt, the agent began performing tens or hundreds of searches for each meeting. Each additional specialized agent can create another recurring source of demand.
He does not believe agentic queries have overtaken human queries yet. Some people operating across many agents may already be generating 100 to 1,000 times their former search volume, he said, but they are outliers. The transition remains early.
The eventual shift he expects is from pull to push. Today, an agent is usually asked to find something now, either directly by a user or by another agent. In the next stage, users may define conditions under which the web should trigger work: notify an agent when a company changes its customer commentary, when a relevant event becomes visible in satellite imagery, or when another agent produces a result worth acting on.
For Parallel, that turns the web from a source consulted at a point in time into a continuously changing feed. The system’s job becomes not only finding information on request, but detecting changes that should cause an agent or a person to act.
The parallel web is published for two audiences
The name Parallel reflects Agrawal’s expectation that web publishing will increasingly serve two audiences. Pages will still need to work for people, but they will also need to be legible and useful to agents.
The company was initially incorporated as Shapley Inc., after the attribution framework that still informs its proposed content market. Agrawal said the name was never intended to last; he considered it unsuitable for a B2B product. Parallel eventually fit both the company’s technical work and its broader metaphor of a parallel web built for AI systems.
Some high-value material may already be agent-first in consumption. Reed pointed to earnings transcripts, which may be more often consumed through agents than by people listening to calls or reading full transcripts. Agrawal made the same point about Parallel’s API documentation. Its customers build AI systems with AI assistance, he said, so agents reading the company’s documentation and SDK code are often the primary audience. Parallel tests its documentation accordingly.
That dual-audience shift connects the company’s technical and economic bets. If agents become persistent users of the web, search systems will need to retrieve evidence for them rather than optimize for clicks. And if content owners are expected to remain open to those systems, Agrawal argues, the web will need a payment mechanism that recognizes the value their material contributes.




