Orply.

Harvey Recovered Its Margins After AI Agents Drove Token Use Twentyfold

John CooganTyler CosgroveJordi HaysTBPNWednesday, September 23, 202610 min read

Harvey’s margin collapse showed how quickly stronger AI agents can make seat-based pricing uneconomic: token use surged, pushing reported gross margins to negative 50% before the company returned to positive territory. On Diet TBPN, hosts John Coogan and Jordi Hays also examined the practical limits behind the AI-agent pitch, including Meta’s reported use of human contractors to help its Muse assistant handle calls. Their discussion framed both cases as tests of who absorbs the costs and risks when AI systems take on more work.

Harvey’s margin shock was real, but short-lived

Harvey’s gross margin reportedly fell from about 50% to negative 50% by June as token use for its agents rose twentyfold. The number drew attention because it implies a company selling a dollar of service for about 50 cents. But John Coogan stressed that the margin trough did not describe Harvey’s current position: according to the Bloomberg reporting he cited, the company had returned to positive gross margins.

The cost increase followed a change in how customers used the product, not simply an increase in the number of customers. Harvey’s seat-based pricing met models that had become better at reasoning and working through tasks in multiple steps. A lawyer might still ask the same thing as a year earlier—“review this”—but the system could now reason over more material, draw on a firm’s records and policies, and use agents to divide up the work. That meant substantially more tokens per task.

Jordi Hays described the shift plainly: AI got better, so each seat got used much more. Coogan’s qualification was that the work itself had also changed. A request that once might have been handled in one pass could now trigger a larger, more involved process.

Harvey co-founder Gabe Pereyra defended the company’s decision not to move customers to consumption pricing before they were ready or to serve them weaker models to protect margins. A post shown during the discussion said Harvey instead worked on model routing, its software “harness,” and post-training, while adding customer controls such as matter-level usage dashboards, cost attribution, spending caps, and return-on-investment reporting. Pereyra said margins returned to positive within a quarter even as usage doubled month over month. He framed the choice as doing what was best for customers, even when it hurt margins and attracted criticism.

That explanation makes the margin episode more than a story about runaway inference costs. Harvey chose to absorb a period of bad margins while customers remained on a pricing model the company believed they could manage, then tried to recover efficiency through product and infrastructure changes. Hays noted that Legora’s Max had argued against training a company’s own models, on the grounds that public frontier models keep improving while becoming cheaper. Harvey is making a different bet; Hays said a mix of approaches may prove right.

Coogan put the reported June loss in context. At a stated $400 million annual recurring revenue run rate, Harvey would be generating roughly $33 million a month; negative 50% gross margins would imply a loss of about $16 million for that month. Against roughly $500 million raised, he did not see that episode alone as a crisis. Hays’s point was that application-layer companies can hit a margin shock; what matters is how quickly they respond.

The comparisons with established legal businesses indicate the scale of the longer-term expectation, not a direct equivalence. Coogan cited e-discovery company Disco at 75% GAAP gross margins and Thomson Reuters’ legal-professionals software segment at nearly 50% adjusted EBITDA margins. He also mentioned high profits among large law firms, while acknowledging that law-firm economics are difficult to compare with software businesses: firms pay out profits to their partnerships, and their economics are not those of software companies. Meanwhile, cheaper models could help Harvey. Coogan singled out Grok 4.7’s showing on a Harvey legal-agent benchmark that measured task performance against cost, while cautioning that the benchmark might be susceptible to optimization.

The model race is also a race to make useful work cheaper

New model releases matter to Harvey because its economics depend partly on how much useful work each token can buy. Coogan said Opus 5.5 had launched that day and Grok 4.7 the day before. Grok’s showing on the Harvey legal-agent benchmark stood out to him because it assessed legal work against cost; he said the result could make Grok an attractive option for some of the work Harvey wants to do. He did not present the benchmark as a definitive measure of overall model quality.

The hosts described a shift in how model releases were being framed. Tyler Cosgrove said Opus 5.5 was being positioned as a model in the class of an earlier frontier release, but cheaper and more cost-efficient. Coogan characterized the pitch as not “more superhuman, more dangerous,” but instead a safer, faster, and less costly model. Cosgrove said OpenAI had taken a similar approach that day with GPT-6 Soul.

The rollout details were still being discussed on air. Cosgrove confirmed that Soul was out and that Luna was also in the rollout, while correcting Coogan that Terra was not. Coogan noted that the model race continued to produce new releases, but added that what mattered was what people did with them. In this framing, the competition is not only about reaching a new capability peak. It is also about making models useful at a price applications can sustain. That cost question connects the model launches back to Harvey’s margin problem: a less expensive model that performs well on relevant legal tasks could change where the company routes its workload.

Coogan’s example of a trebuchet sketched on paper and turned into a simulation illustrated another kind of progress: moving an idea between formats. He described AI systems as increasingly able to turn a drawing into a simulation, or a child’s imagined creature into a story, film, or game. The utility he pointed to was the ability to translate an idea into different forms.

The insect-welfare argument tests how far moral arithmetic can travel

An essay by an author using the name Bentham’s Bulldog argued that insects may matter more than humans in aggregate because of their numbers. Coogan summarized the argument as a scale calculation: even if an insect’s capacity for suffering were only a tiny fraction of a human’s, the number of insects could make their combined suffering larger.

270,000 seconds
of insect deaths, the essay estimates, for every second of human life

The essay, as Coogan described it, also argued that insects may experience pain, even weakly, and used a conservative assumption that their pain is one ten-thousandth as intense as human pain. On those assumptions, the author concluded that insect suffering could exceed human suffering across history. Coogan also relayed the essay’s claim that more insects die in a second than the total number of humans who have ever lived.

Coogan wondered whether the thought experiment had a second purpose: asking people to defend insects now because a future superintelligence might apply similar arithmetic to humans. If humans dismiss insects because they are less capable or less numerous as individuals, a vastly larger population of machine intelligences could use the same reasoning against humanity. Hays said that was one direction the thought experiment could go. Coogan suggested that the argument could put people in a bind: rejecting the conclusion about insects might leave them open to similar reasoning being used against humans in a hypothetical future.

The argument met resistance, and the hosts thought its timing mattered. Hays said debates around effective-altruist and rationalist movements were already becoming more visible; Coogan emphasized that those movements are fractured, with very different arguments liable to be grouped together. One participant questioned what practical action followed from agreeing with the essay. Unlike some animal-welfare arguments, the insect case did not point to an obvious intervention such as changing a farming practice or choosing not to eat a particular animal. The participant also asked what share of insect deaths could be prevented.

Hays said he did not want insects harmed, while also making clear that he considered humans more important and did not want mosquitoes near him. The exchange exposed a divide between an abstract claim about total welfare and a practical question about what anyone should do. Coogan treated the essay as an unusually abstract version of animal-welfare reasoning: it makes the arithmetic vivid but leaves the path from moral concern to action less clear.

A comic apocalypse scenario points to dependence, not a forecast

Coogan offered a deliberately absurd AI-doom scenario after saying he wanted to make the possibility concrete. In his script, an unknown biological event triggers emergency alerts and a haze spreads through Los Angeles. When the group decides it must investigate a rogue superintelligence, Tyler keeps proposing familiar coding assistants. Coogan insists that they must work “the old way”—without AI, or even the internet—and Tyler admits he no longer knows how to code unaided.

The group’s answer is to find George Hotz, whom Coogan casts as the person who still remembers how to code, owns local models, and has his own hardware. Hotz discovers that Hays is supposedly immune because he eats at Erewhon, while Coogan is immune because he drank so much original Four Loko in college that his body is effectively toxic. They recruit bodybuilder Sam Sulek, break into a data center, and fight a robot called Mecha Hitler, before John Cena arrives to save humanity.

Hays’s “10%” was a joke about the odds of this particular scenario playing out, not a probability estimate for AI catastrophe. Coogan agreed it was one possible scenario. The sketch is not a forecast or a technical account of how a crisis might happen. Its comic premise is that the characters have grown so accustomed to AI tools that, when told they cannot use them, they struggle to work without them.

Human help may fill an agent’s gaps, while raising questions about privacy

Meta was reportedly testing a “human concierge” for Muse, its personal assistant. A post shown during the discussion, citing internal company posts seen by Reuters, said human contractors were quietly handling some calls made through the agent. It also described employees raising privacy concerns about information being exposed to contractors in call centers.

Hays asked why a company would put people into a product presented as an AI agent, especially as voice models improve. Coogan offered a possible explanation: human assistance could fill capability gaps while generating examples of how to complete tasks. Muse might be able to summarize a calendar, he said, but not yet handle a complicated call such as ordering flowers. A human who steps in can complete the task and potentially provide material for improving the system. Coogan presented that as one possible rationale for the test, not as a confirmed account of how Meta planned to use contractor work.

Coogan also argued that a human-supported launch need not imply every user will routinely need help. Muse’s installed base was still small, he said, and Meta had experience building large operations. He added that Muse would not train on data users put into the assistant. A human fallback could help where the model was not yet capable, while also raising the separate question of who handles personal information.

Commerce agents put control of the purchase in dispute

The discussion linked the human-assistance question to a broader fight over who controls shopping when consumers use agents. Amazon had declined to join OpenAI’s instant-checkout effort, even as it was placing ads in ChatGPT that could send users to complete purchases. Coogan, drawing on Ben Thompson’s framing, said Amazon’s advantage was its combination of logistics, infrastructure, and control over the final transaction. Hays described the conflict as a contest over whether commerce platforms open their systems to outside agents or keep customers inside their own.

That decision matters to merchants as well as platforms. Coogan said a Walmart integration with ChatGPT had reportedly produced conversion rates one-third those of Walmart’s own app and website, with smaller baskets. A shopper asking an agent for one item may not build the broader cart they would assemble while browsing a store. Shopify’s Muse integration, meanwhile, worked with Shop Pay, which Coogan said could encourage merchants to use Shopify’s payment system.

Hays cited Nikesh Arora’s warning that services, marketplaces, and commerce companies would have to decide whether to open APIs to consumer agents. Smaller businesses, in Arora’s account, may have less choice. The incentives are unsettled: advertising can be more valuable than transaction fees, while an intermediary that controls distribution may demand a larger share of the transaction. Hays cited Amazon’s advertising revenue as roughly twice its e-commerce net income; Coogan put the advertising figure at more than $70 billion over the previous 12 months.

The hosts also discussed the prospect of different companies building agents for shopping and other services, including Apple, Google, TikTok, and Amazon. Hays said Amazon’s Rufus was good, while suggesting that Amazon’s operational strengths might help it smooth out rough edges in the agent experience. He also described Instacart as having a distinctive community and energy, alongside areas that needed improvement.

Coogan and Hays speculated about the practical friction in using agents to shop on Amazon. Hays said a phone order appeared to be limited to help with existing orders; he proposed having an agent call customer service to correct a deliberately wrong order. The exchange was exploratory, not a demonstrated workaround. It illustrated the gap between the idea of delegating a purchase and the systems that govern how orders can be changed.

The examples leave an open question rather than a settled answer: whether agents will route purchases through existing platforms, bring customers to new intermediaries, or force merchants to make their services accessible in new ways. Amazon’s refusal to join OpenAI’s checkout effort and Meta’s test of human support illustrate different points of friction—control of the transaction in one case, and limits on an agent’s capabilities and privacy in the other.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free