Orply.

AI Excels at Verifiable Work, While Building Still Depends on Human Judgment

John CooganJordi HaysTBPNThursday, October 8, 202610 min read

OpenAI’s mathematical results illustrate the rapid progress John Coogan sees in tasks that can be carried out and checked in digital environments, but he argues that technical capability does not settle the broader work of building a company. That work still involves judgment, hiring, dealmaking and execution. In the discussion of SpaceX’s proposed $40 billion chip financing, Jordi Hays argued that investors must assess the risks and valuation themselves, even if Elon Musk’s track record earns him some latitude.

The film review was a joke; the cinematography was not

Jordi Hays opened with a deliberately useless review of Primetime, the Robert Pattinson film directed by Lance Oppenheim: the audio was good, the video was good, and all three acts were good. The joke gave way to a more specific reaction. Hays and the other hosts praised the film’s 4:3 frame, its grain and blown-out highlights, and the way its 23 cameras seemed to evoke different television cameras from the period.

Hays called the film a creative throwback to how people watched television at the time. He also described its treatment of Chris Hansen as a heavy-handed but still interesting investigation of Hansen’s career and how he is remembered: the film seemed to make its judgment clear without entirely preventing viewers from making their own. Hays said Pattinson was “in flow state”; John Coogan agreed.

The show paired the review with its Christmas countdown: 78 days remained. Hays pointed viewers to an earlier interview about the Christmas-tree industry, including how it is structured and how to keep a tree alive, and encouraged donations to Trees for Troops. He said he had donated and noted that the organization accepts company-matched gifts. He also recounted the interviewee’s joke about plastic trees, calling them “toilet brushes,” while acknowledging that a reusable plastic tree may make more sense for people who cannot afford a new real tree each year.

AI advances fastest where work can be simulated and checked

Coogan described two different rates of progress: tasks that take place entirely in a virtual environment are accelerating quickly, while work that touches the physical world, depends on human-generated data, or needs a longer-running flywheel moves more slowly. He used that contrast to connect OpenAI’s mathematical results with a set of open-source Adobe lookalikes. In both cases, capability appears to move fastest where a system can operate and be evaluated within a bounded digital environment.

The Artcraft project presented free, open-source tools as reimplementations of Adobe products: Photocraft for Photoshop, Vectorcraft for Illustrator, Filmcraft for Premiere Pro, Lightcraft for Lightroom, Effectcraft for After Effects, and Designcraft for InDesign. Coogan said he had tried Photocraft. It was “not bad,” and reproduced a surprising number of tools, though it lacked some features he expected. The project described the products as “reimplemented from scratch”; Coogan said he did not know whether every Adobe product had been decompiled.

For him, reproducing a desktop interface did not reproduce every service behind it. Photoshop’s older fill tools worked from nearby pixels and textures; its newer generative fill uses server-side models. That feature did not work in the clone he tried. Coogan said Photocraft was not a perfect substitute for demanding professional work. But for a smaller job—compositing two images for a meme or adding text—a prompt might replace twenty minutes in Photoshop with a minute in ChatGPT or another tool.

Hays described the change from the user’s side: someone who once needed time or help to make an edit can now ask for it in a prompt. Coogan agreed, while distinguishing occasional edits from spending a full day in After Effects or Premiere Pro. A prompt could remove friction from the former; he did not think it yet replaced the sustained workflow of the latter. He also said server-side, inference-heavy features give software makers something of a moat, even as that moat shrinks.

That unevenness helps explain why OpenAI’s math announcement seemed consequential to Coogan. OpenAI said an internal frontier model had produced a broad range of new mathematical results and that it had consulted an independent advisory group at the Institute for Advanced Study. The on-screen catalog listed 722 manuscripts across 372 result families and said verification varied by manuscript. Hays noted that the manuscript count did not straightforwardly establish how many distinct problems had been solved. Coogan likewise observed that OpenAI’s announcement called the work a “broad range,” rather than offering a simple total of problems.

722
manuscripts across 372 result families; verification varies by manuscript

Coogan cited François Chollet, creator of ARC-AGI, on one possible explanation for the uneven progress. Math and code, Chollet suggested, may be pushed much further through reinforcement learning with verifiable rewards, while other areas may plateau sooner because they remain bottlenecked by human-generated data. The open question is whether gains in less-verifiable areas are a side effect of broader improvements in general intelligence, or mainly reflect the continued addition of human data. Coogan said much depends on which explanation is right.

For Coogan, the results felt less like a familiar AGI milestone than a “superintelligence moment”: he said no human had solved this many math problems at this level, so quickly. But he also stressed that the capability is “spikey.” A system can excel where work can be trained and checked without fitting neatly into a general human framework. He preferred “machine intelligence” partly because its capabilities have a different shape.

Making software more malleable does not remove the systems around it

The same divide between what is easy to reproduce and what remains difficult to change came up in Coogan’s discussion of macOS. He relayed Ben Thompson’s frustration with permissions when building agentic software: an app created by an agent may need access to files or the network, and macOS requires the user to grant permissions. As Coogan described Thompson’s proposal, a coding harness would create software that operates with the access and permissions already granted to the agent.

Coogan suggested that AI could make Linux easier to use by letting people prompt their way through tasks that might otherwise require command-line work, such as changing a setting or fixing a driver. He also wondered whether AI-assisted decompilation might eventually let people modify macOS itself, including features they dislike, while keeping the rest of the Apple ecosystem. Another participant compared the possibility to jailbreaking and noted that modifying a system can mean losing features. Coogan wondered whether a modified installation could still receive Apple services and security updates; he also pointed to Apple’s hardware advantage.

The broader question was not only whether software can be copied, but whether the surrounding services and permissions can be carried over. Coogan said the Adobe clone was early-adopter territory and did not mean Creative Cloud users were cancelling subscriptions en masse. He cited Adobe’s share-price declines as part of a wider “SaaS-pocalypse” narrative, while saying he had not seen signs of mass cancellations. The clearest near-term pressure, in his view, was on smaller, less professional tasks where a prompt could replace the effort of opening a full application.

There was a parallel uncertainty in how people respond to machine-generated knowledge. Hays read a post by Roon urging people not to become spectators who applaud discoveries without understanding them. Coogan pushed back on the idea that being impressed without dissecting a result was a failure. He likes watching a magic trick without needing to know how it works, he said; the math results impressed him, though he did not feel obliged to read every paper.

Technical ability is not the same as the work of building

The hosts connected the math results to a warning from Ilya Sutskever: people who value intelligence above every other human quality are “gonna have a bad time.” Coogan described the show’s “golden retriever mindset” as being friendly, physically fit, avoiding unnecessary enemies, and not assuming you are the smartest person in the room. He was not arguing that intelligence no longer matters. His point was that it is only one part of the work.

Coogan contrasted high-level mathematical performance with the day-to-day work of building a company. He noted the strong pipeline from International Math Olympiad achievement to startup founding, but said the work of someone such as Cognition’s Scott Wu does not look like solving Olympiad problems. It involves dealmaking, hiring, strategy, implementation, and product development. The joke that a company could grow by hiring mathematicians to solve a fixed number of billion-dollar problems each year underscored the gap: that is not generally what the work consists of.

A related argument appeared in a post about quantitative finance: quant firms were hiring “idea guys,” including thinkers, liberal-arts graduates, and generalists, not only quants. Coogan said he knew investors with backgrounds in philosophy and history who brought a distinctive lens to their work. He mentioned people drawing on years of historical study to think through how the Eurozone might develop.

The hosts also discussed a letter from AQR asking what remains a durable differentiator once AI has surpassed many human analytical and technical capabilities. The letter’s answer was creativity and emotional intelligence—not because these qualities are wholly immune to AI, but because people can develop an advantage in them. Coogan connected that argument to the hiring post, but the discussion presented it as a view about what may matter, not as proof that quant firms generally have changed their hiring practices.

A $40 billion chip order puts the burden on investors to understand the bet

Jordi Hays described SpaceX as seeking about $40 billion to finance a large Nvidia chip order: roughly $10 billion in bank loans and $30 billion in investment-grade debt. Apollo was expected to lead the financing effort, with Pimco among lenders in talks. Hays said SpaceX’s BBB rating would make the debt available to some insurance and pension funds that are more limited in their ability to buy junk-rated notes.

The plan would deepen SpaceX’s ties to Nvidia as Musk’s companies pursue AI projects. Hays read a statement saying SpaceX had chosen to build exclusively on Nvidia in the near term because it considered the Vera Rubin architecture the best AI computer and valued the partnership. John Coogan asked how that fit with Musk’s work on Dojo and in-house chips at Tesla. Hays’s interpretation was that Nvidia was the strategic choice for now: “kiss the ring.”

The financing memo became a test of how much explanation investors should expect. Investors approached about the purchase reportedly received a two-page document with pictures of outer space and an arrow indicating that SpaceX planned to build data centers somewhere in the universe. One investor asked how to take that to an investment committee. Coogan found the brevity funny, but also said the basic proposition—spend $40 billion on chips—could be sketched in simple terms. The harder questions are how to value the chips and how the plan could fail.

Hays said investors considering billions for AI infrastructure should do substantial work of their own on those risks. He also argued that Musk had earned some latitude because trusting his process had worked to date. That might justify a shorter-than-average memo, Hays said, but leaves investors responsible for understanding the bet.

The infrastructure discussion also touched on public reaction to what gets built and what jobs remain. Coogan praised a photograph circulated as a European data center because the building looked more like a traditional station than an industrial facility. Hays said it was a French train station, not a data center, and noted that the image had been Community Noted. The hosts did not settle whether the original post was a joke. Coogan nevertheless took its circulation as evidence that people respond to the idea of more attractive infrastructure.

On construction work, the hosts played a reporter asking a worker at a Pennsylvania data-center site what happens after the build-out, when people say the jobs are temporary. The worker replied, “Every job in my field’s temporary.” Coogan said the answer made a straightforward point: a construction job ending with a project is not, by itself, unusual for the field.

A car window could become ad inventory—but the economics are still a guess

The hosts ended on HoloGlass, a product from Pranos that Coogan described as a way to turn a car window into a display. The demonstration showed a projector mounted inside a vehicle and content visible through its side window. Coogan said the system could display videos or custom signs generated through an app, and suggested that drivers might use it to run ads while commuting. Hays asked whether a driver could show something like Paw Patrol to people stuck in traffic; Coogan said the user could display whatever they wanted.

Coogan called the idea “extremely cyberpunk” and joked about adding ads to a Cybertruck or Jaguar. He also pointed out that the hardware looked like an off-the-shelf projector paired with a mount and software. That did not diminish his enthusiasm: he said he had been worried there would not be many more advertising surfaces to add.

The hosts did not have a firm estimate of what a driver could earn. Coogan said the product started at about $300 and guessed, without claiming to know, that a driver in Los Angeles might make around a dollar a day. He said the audience and location would matter: ads aimed at venture capitalists around Sand Hill Road could have a different value from ads shown in places where people have less disposable income. Hays thought the economics might work if the device could be paid off in six months, after which the driver would effectively be paid to commute. That was a possibility they were entertaining, not a revenue estimate supplied by the company.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free