Orply.

Cambridge Proposal Calls for Comparable AI R&D Metrics and Crisis Plans

John CooganJordi HaysTBPNTuesday, September 29, 20269 min read

Government contact with AI lab leaders is increasing as Cambridge researchers call for comparable measures of AI’s role in research and development and plans for severe failures. TBPN hosts John Coogan and Jordi Hays also raise a specific market risk: if widely used agents make similar financial decisions at the same time, they could contribute to a bank run or flash crash. Coogan cautions that adoption and trust will take time.

AI safety has moved from public rhetoric to government contact

John Coogan reads the Saturday Night Live parody of Anthropic CEO Dario Amodei as a sign that AI has entered mainstream political culture. The joke worked, he said, because Amodei had become recognizable through appearances on major television networks. Jordi Hays pointed to the sweater vest as part of the character: Amodei’s manner and appearance were familiar enough to be imitated.

Coogan called the sketch a possible “crossing the chasm” moment, while qualifying that Michael Che seemed to stumble over Amodei’s name during the introduction. Perhaps, Coogan suggested, the name was not yet familiar to the writers. Either way, the parody placed a private-company lab leader in a public argument about who should protect people from AI—and whether that rhetoric is itself becoming a target.

The hosts also treated public presentation as part of the politics. Hays argued that Amodei could have more impact by not smiling while discussing catastrophic risks. Coogan compared the problem to the austere public manner of officials who deal with consequential subjects: a serious performance signals that the speaker takes the matter seriously. He contrasted that with Mark Zuckerberg’s more upbeat approach, which he said may appeal to people looking for a more optimistic posture, but not necessarily to researchers working at the frontier. Hays agreed that Zuckerberg’s enthusiasm for AI may not persuade those researchers.

That public profile now sits alongside direct political engagement. Amodei met President Donald Trump for a private dinner, after being notably absent from a dinner that included other technology leaders and China’s Xi Jinping, according to the hosts’ discussion. Coogan described the meeting as part of a faster cadence of contact between government and AI companies: meetings were happening “every week now, basically.” The hosts did not know what would emerge from the dinner, but treated the combination of political meetings, public debate, and cultural attention as evidence that lab leaders could no longer stay outside the conversation.

Cambridge’s proposal asks for measurement and contingency plans

The University of Cambridge proposal discussed by Coogan went beyond the familiar call to “talk about” AI safety. Its first concrete recommendation was to establish measures of how much AI is contributing to research and development at leading labs: what share of the code, spending, or other R&D work is being done by AI.

Coogan said labs such as Anthropic and OpenAI had already described AI’s role in their research, but their public measures were difficult to compare. One might describe a model as operating at an intern level; another might point to a model generating a thousand ideas. Those are different claims, he said, and there was no clear apples-to-apples way to judge which lab was further along. The proposal, in his telling, was to develop consistent metrics so that progress in AI-enabled R&D could be monitored rather than left to competing descriptions.

The second major thread was government preparation for severe failures. The proposal called for plans to respond to hypothetical scenarios such as a cyberattack in which a rogue agent brings down parts of the internet or hacks into systems, as well as biosecurity threats and economic disruption. Coogan compared the need to disaster response: governments have plans for hurricanes, including people, food, and shelter, but he said there was no equivalent framework for an AI-related event that took down the internet. The researchers, as he summarized them, were identifying where preparedness was needed, not presenting a full set of solutions.

Hays joked about creating disaster simulations where participants try to shut down a rogue data center. The joke pointed to a serious gap in the proposal as discussed: identifying scenarios and measuring capabilities are steps toward preparedness, but the hosts did not describe a settled operational response.

Some safety concerns are shaping personal decisions, but a bunker can make you a target

A Wall Street Journal article reported that some early Anthropic employees were considering buying remote land in the United States in case AI went seriously wrong. The article also described employees trading disaster scenarios and preparations in a private Slack, and mentioned ideas circulating in the wider safety community, including iodine pills, remote islands, and shielded desert bunkers.

The hosts saw the land purchases as a concrete answer to a challenge often put to people who fear catastrophic AI outcomes: if the risk is real, why does it not affect your financial choices? But they immediately questioned whether buying a remote retreat would actually help. A person who arrives in a community they do not know with a well-stocked bunker may have built what preppers call a “loot drop”: a cache of supplies that others could take in a crisis.

That concern complicates the familiar image of individual preparedness. Survival would depend not only on the bunker’s location and supplies, the hosts argued, but on whether its owner was part of a local community. Coogan’s hypothetical was a person who had built a remote shelter without meaningful ties to the people living nearby. In a breakdown, the shelter could become a target rather than a refuge. The hosts suggested that being prepared as a group, with people already connected to one another, might reduce the risk of conflict over resources.

Their discussion moved between speculative scenarios and the difficulty of planning for them. Coogan imagined the prospect of having to defend oneself against rogue machines, then said he would hate being in that situation. The practical point was less about any particular scenario than about the social assumptions behind preparedness: a private stockpile may not protect its owner if the surrounding community is excluded from the plan.

Instinct’s growth points to a new market for agents—and a trust problem

The hosts turned to Instinct, a personal-agent company whose founder, Noah Shinn, had spoken with investor and interviewer Patrick O’Shaughnessy. The figures they discussed came from O’Shaughnessy’s posted highlights. One headline figure—$1 billion in transaction volume—was ambiguous: John Coogan said it might refer to annualized volume or total volume to date. He compared it with Stripe, which he said took about 18 months to reach $1 billion in annualized volume, after spending two years building in beta before launch. Coogan said Instinct reached $1 billion in transaction volume in closer to six months, but the uncertainty about what the figure measures makes the comparison provisional.

10%+ a day
reported Instinct growth, with zero dollars spent on marketing

Other figures in the highlights were presented as reported growth metrics, not as clarifications of that $1 billion volume claim. Compute demand was said to be doubling roughly every week. Forty percent of users reportedly shared a personal credit card with Instinct within three weeks, and users who connected at least one piece of sensitive information retained at about 80%. The highlights also said some users sent more than 90% of their messages by voice, and some small businesses ran their entire back office on the service.

For Jordi Hays, the product’s appeal was partly that most people outside technology had not yet experienced an AI system doing work for them, such as making a product. The novelty could help explain why users shared screenshots and why access remained invite-only. Coogan wondered whether even screenshots of mistakes might attract curious users; Hays thought they could.

The interview highlights also described different compute setups for background work as three, five, or eight times more efficient on the same compute. They said buying compute at the last minute cost three to four times more, and that Shinn spent about 40% of his time on compute. Those details help explain why compute planning is central to the company’s operating model, alongside the reported weekly doubling in demand.

The commercial model remains a question. About half of Instinct’s transaction volume was reportedly travel. Shinn’s stated goal, as Hays said, was to keep the product free for life, while the company planned to take a cut of what users bought. The interview highlights compared possible take rates with Shopify’s 2.5–3%, Amazon’s 15%, and Apple’s 30%. They also suggested that agents might reduce advertising revenue while increasing purchases by making buying easier. Hays pushed back on the premise that checkout itself was still much friction: on many Shopify stores, a customer who knows what they want can already buy quickly. Coogan’s counterexample was the less considered purchase—finding a workable set of cable-management products, for instance, where a user does not care much about brands and would rather delegate the search.

Instinct’s transaction volume does not yet mean it earns a share of those transactions. Coogan said purchases currently flowed through third-party cards, so the company was not making money from them. He and Hays discussed the possibility of an Instinct-branded credit card, which could create interchange revenue without asking users to opt in to a separate fee. Coogan said even a small margin on a large volume could become meaningful, particularly for a team of 14 people. He pointed to the company’s reported $1 billion fundraise at a $10 billion valuation, less than a month after a $2.5 billion valuation round, as evidence of investor confidence in its growth and monetization possibilities.

But that confidence depends on continued adoption and on the product holding its position as competitors develop. The hosts also noted that the company was less than a year old, that invites had reportedly sold on eBay for around $300, and that Shinn was 23. They treated these as signs of unusual early interest, not proof that current growth would persist.

Agents could coordinate without communicating, but adoption will take time

A separate concern raised by the hosts was that widely used personal-finance agents could make similar decisions at the same time. A post they discussed cited warnings from Apollo’s chief economist and former SEC chair Gary Gensler that algorithmic decisions made at scale could undermine market stability. If many agents independently concluded that the same move was optimal, the result could be a bank run or a flash crash, even without the agents communicating or colluding.

Hays framed the broader mechanism as the removal of friction. An always-on agent might notice that a user had not touched a product or service in a long time and ask whether to keep paying for it. The same pattern could apply to money: an agent might move funds into a higher-yield account or another investment. Hays argued that this kind of proactive service already exists in good banking experiences, though many people may not have access to it.

Coogan agreed that agents could eventually act on users’ behalf, but argued that the transition would not be immediate. Adoption was still limited, he said, and users would need to authenticate, link accounts, and learn to trust agents with financial decisions. Early systems might ask permission before acting; later, users might trust them to act without asking. Building that relationship, in his view, could take years.

He compared predictions that agents would rapidly displace existing services to early claims that ChatGPT would make Google search revenue collapse. The replacement may be plausible, he said, but diffusion takes time—and institutions can adapt during that period. The potential for coordinated decisions remains, but the hosts differed in emphasis: Hays stressed that agents could remove friction and make similar actions widespread, while Coogan argued that adoption and trust would take time.

Meta is packaging its AI stack for business customers

Meta announced a new Enterprise Platform and named former MongoDB CEO Chirantan “CJ” Desai as its Chief Enterprise Platform Officer, reporting directly to Zuckerberg. The announcement said the offering would bring Meta’s models, agents, infrastructure, and business tools to companies and developers. The named products included the Muse agent, Meta Business Agent, Muse API, and Muse Code.

Hays interpreted the announcement as the creation of a new business unit, but was skeptical that its eventual revenue would be easy to parse. He predicted that Meta could make a large compute deal and then report rapid growth for the enterprise platform, potentially from zero to $10 billion in annual revenue. That was Hays’s forecast, not a figure in Meta’s announcement.

Coogan noted that enterprise-platform revenue could come from several sources: tokens, inference, compute sales, bare-metal deals, or services. The hosts left open whether Meta’s new unit would represent a distinct business or bundle together different kinds of AI and infrastructure revenue.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free