AI Capability Is Moving From Benchmarks to Completed Tasks
John Coogan and Jordi Hays argue that the important tests for new technology are operational rather than promotional: GPT-6 Astra should be judged by whether its apparent Blender competence carries into real production work, and model costs by completed tasks rather than token prices. They apply the same standard to Tesla’s Cybercab, which they say will matter only if Tesla delivers a sub-$30,000, steering-wheel-free vehicle capable of viable autonomous operation by the end of 2026.

Blender makes agentic capability easier to recognize—and raises the question of transfer
John Coogan found GPT-6 Astra’s most consequential evidence not in its reported benchmark results but in demonstrations of the model working inside Blender. The outputs shown included a dusk render of San Francisco’s Palace of Fine Arts, a detailed modern-villa scene, and a demo house carried from Blender into a walkable Unreal Engine 5 environment.
The Palace of Fine Arts example, posted by Sharif Shameem, recreated the landmark as a lit 3D scene with water, a rotunda, colonnades, and sculptural detail. Coogan’s interest was not that the model had somehow created the building without available reference material. In his view, a reconstruction could potentially draw on internet images, diagrams, architectural plans, blueprints, or existing online 3D models. What appeared notable was Astra’s ability to work through those possible inputs and carry out a modeling task inside the software.
Jordi Hays noted that Matt Shumer, who had shared some of the strongest examples, had been saying that “the art of prompting is back.” The most impressive results may require more than a single sentence: a manager agent, sub-agents, and iterative refinement. But that does not change the workflow shift Coogan was reacting to. The system appeared able to operate in a production environment rather than merely generate an image in a chat interface.
A side-by-side Blender villa comparison shared by Hesam made the point visually. Fable 5.1’s output was shown as a simple, low-poly scene; Astra’s included furnished interiors, a pool, detailed materials, and more realistic lighting. Hesam contrasted that gap with the Artificial Analysis intelligence scores displayed in the post, where Fable scored 66 and Astra 61, and said something must be wrong with a measurement that produces that ordering.
The demo house from Thomas Ricouard suggested a larger production chain. Ricouard described using Astra to build a house in Blender and bring it into Unreal Engine 5 as a walkable experience. Coogan called the result properly lit and thoughtfully designed, while leaving open how much came from prompting, source material, or the model’s execution.
Everything becomes translatable. You'll be able to take your tweets, turn them into blog posts, turn them into videos, turn them into 3D games on Steam very soon.
For Coogan, 3D modeling is a more intuitive capability test than an abstract score because he knows the underlying work. He has spent hundreds, perhaps thousands, of hours in Cinema 4D and Houdini. Building a scene means learning unfamiliar functions, adjusting settings, arranging objects in a three-dimensional space, choreographing a camera over time, and waiting for previews, caches, and simulations. A smoke effect, fracture, or animation can require substantial computation only to reveal that an earlier setting needs to change.
That friction made even a small custom visual expensive. While publishing weekly YouTube videos, Coogan commissioned a 3D battleship board for an analysis of U.S.-China military competition. The work could occupy an editor for a week. An agent that materially shortens that cycle could make a greater volume of visual work economically feasible.
Real-estate rendering was his immediate commercial example. Zillow already lets users restyle photographs for some listings. Coogan suggested that a system capable of producing or navigating a 3D representation of a home could lower visualization costs, particularly for spec homes that cannot justify conventional rendering budgets. He and Hays saw a Jevons-paradox possibility: cheaper rendering may lead to more rendered spaces rather than simply eliminating the people who make them.
The harder question is whether Astra’s competence is specific to Blender. Coogan called Blender an unusually favorable environment for this kind of training because it is open source and can be downloaded and replicated widely in reinforcement-learning environments. He also referred to reports that OpenAI had bought “tens of thousands” of Mac Minis and Mac Studios for computer-use training.
The same dynamic could favor software that happens to be used inside frontier labs. Coogan pointed to Slack, which he said Meta’s Alex Wang had described as better for working with AI agents. His theory was that frontier labs use Slack themselves, making it a natural training environment; models may become especially capable in it; and that competence could in turn make Slack more appealing to organizations deploying agents.
That creates a new form of switching pressure. Coogan had seen Blender’s ecosystem gaining momentum by 2021, but still faced the accumulated friction of leaving Cinema 4D: plugins, textures, rendering settings, and learned interface habits. Models may eventually give users another reason to move from Cinema 4D to Blender if they work better in Blender. Or they may generalize from Blender to adjacent packages such as Houdini and Cinema 4D.
The distinction matters because the tools are not interchangeable. Coogan noted that different programs remain better suited to different work, including smoke and water simulations. Hays said there were reports that Astra was effective at video editing, but Coogan said most video is still edited in Adobe Premiere. Its closed nature, licensing terms, and access restrictions could make large-scale outside training more difficult, he speculated. Open-source editing suites exist, he added, though they lag behind.
Hays did not see creative work following a simple replacement story. More people may attempt renders, interactive environments, or games themselves, then reach the point where the model cannot get to the desired quality and hire someone with expertise.
Anything that can be rendered will be rendered.
Coogan agreed with the direction while stressing that taste still matters. In editing and motion design, the hard part is often not merely following instructions. It is telling the right story, choosing the right visual treatment, and avoiding “slop,” whether the work is made manually or with an agent.
The useful unit of measurement is the completed task
Astra’s launch brought the familiar rush of benchmarks. A comparison shared by Jack Altman put Astra at 98.6% on ARC-AGI-3 and 97.6% on FrontierMath Tier 4, with additional results across agentic, automation, CAD, and software-engineering evaluations.
| Evaluation | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|
| ARC-AGI-3 | 98.6% | 7.8% | — | — | — |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% |
| Agents' Last Exam | 59.3% | 52.7% | 48.7% | 52.7% | — |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% |
| BenchCAD | 95.9% | 83.3% | 84.3% | 87.9% | 82.1% |
| DeepSWE v1.1 | 74.1% | 70.8% | 67.4% | 68.9% | 73.7% |
John Coogan said the numbers do not produce a “feel the AGI” moment for him. He lacks direct intuition for the difficulty of many benchmark tasks and already expects computers to excel at math. New benchmarks, he argued, are quickly hill-climbed: once they are created, models race toward 99% or 100%, while outside observers lose a clear sense of what that performance means.
His preferred evaluation is more personal: point a new model at something one knows deeply. The goal is to avoid what he called Gell-Mann amnesia—being persuaded by apparently fluent output in a domain where one cannot identify mistakes. The hosts joked about horses, comedy, and shrimp fried rice as possible tests. The underlying standard was serious: judge a model on work where the evaluator can recognize competence.
A 3D scene carries that kind of legibility for Coogan. So does a game, although he resisted treating every browser-based generated game as a meaningful achievement. The goalposts, he said, have already moved: a game should be entertaining and substantial, not simply a small experience running in a Chrome tab. He wanted the result to run on Steam, a PC, a PlayStation, or an Xbox.
The same move from proxy to completed work applies to model pricing. Coogan highlighted Steven Heidel’s argument that token pricing is becoming “effectively meaningless.” Heidel’s example compared Gemini 3.8 Flash, which appeared 13 times cheaper per token than Astra, with Astra, which he said could be less expensive per completed task because it uses tokens much more efficiently.
Coogan did not expect aggregate token volumes to necessarily fall. Companies may still use more tokens overall even if individual models become more efficient. But the relationship between token growth and value creation could change: token use may grow more slowly while models complete more valuable work. The relevant comparison, in Heidel’s phrasing, is cost per task rather than cost per token.
The AGI debate becomes more concrete under that standard. The hosts noted declarations from The Wall Street Journal, Stripe, and Tyler Cowen that some form of AGI had arrived or was arriving. Jordi Hays argued that such calls may look reasonable in retrospect because capability develops on a spectrum rather than crossing one universally agreed threshold. Coogan’s remaining question was not whether the systems are artificial or whether they can be intelligent at many tasks. It was how broadly that competence travels.
The August jobs figures put an economy-wide AI apocalypse on hold
The August employment report gave John Coogan a direct counterpoint to claims of immediate, broad AI displacement. The United States added 162,000 jobs, well above the 53,000 gain expected by economists polled by The Wall Street Journal. The unemployment rate held at 4.1%, which Coogan described as historically low and consistent with a generally healthy labor market.
The gains were not confined to one corner of the economy. Food services and drinking places added 59,000 jobs. Local-government education added 42,000 after losses in previous months. Manufacturing and healthcare also registered gains, while information and finance lost jobs—sectors Coogan said are generally regarded as relatively exposed to technology.
June and July were revised upward to gains of 31,000 and 21,000, respectively. July had initially been reported as a loss of 23,000 jobs. Jordi Hays called the upward revision a “narrative violation.”
Coogan emphasized that a handful of high-value technology companies with small employee bases do not describe the broader labor market. He cited Hugging Face as an example: a company with a couple hundred employees and a $12 billion outcome. That is a power-law outlier, he argued, not the basic pattern of how the U.S. economy generates employment.
His conclusion was narrower than a claim that AI will not reshape work: the August figures put an economy-wide AI jobs apocalypse on hold for another month. They also, he said, argued against a Federal Reserve rate cut in September.
Dyson’s innovation does not settle the question of quality
The Dyson debate turned on a distinction between moving a consumer category forward and making products that feel good to own today. A post by TJ Parker calling Dyson products “generally not very good” drew agreement from users who described disappointing vacuums, while Bryce Roberts said his daughters loved their Airwrap.
Jordi Hays made the case for the company before conceding the criticism. Dyson’s business and origin story are “incredible,” he said, and its decision to make household tools look like Halo weapons is a distinctive form of product differentiation. He appreciated the company’s willingness to overengineer mundane objects, even if he did not necessarily endorse buying every resulting premium appliance.
But Hays said his own Dyson vacuums often felt creaky and a little cheap. The futuristic cleaning-device aesthetic, he added, did not fit comfortably into every home. A reply in the discussion raised another practical complaint: strong Dyson suction may make rugs or carpets shed faster.
John Coogan took the historical view. Consumers forget, he argued, how rattly, clanky, cord-bound, bagged vacuums could be. Vacuuming once meant unwinding a cable, finding an outlet, moving the plug room by room, and managing the cord around furniture. He and Hays recalled treating the chore as an optimization problem: identify the fewest outlets needed to reach the entire house.
Dyson, in Coogan’s telling, solved a real category problem. Battery-powered cleaning removed the cord, and the company made the product feel substantially more modern than what preceded it. That does not answer whether current products feel plasticky or underbuilt. It explains why Coogan thought criticism of present-day fit and finish should not erase the scale of the original advance.
Cybercab must be delivered, autonomous, and economically viable
The Cybercab test, as John Coogan and Jordi Hays defined it, is not a preorder, a prototype sighting, or an impressive interior tour. Tesla must deliver the vehicle by the end of 2026 for under $30,000, without a steering wheel, and with enough autonomous capability to perform economically viable tasks.
The hosts were explicit about delivery. A customer needs to pay the stated price and receive the vehicle by December 31, 2026. A reservation does not count. Nor does a supervised-driving system that performs well in ordinary use but cannot operate commercially on its own.
That threshold comes from the promise associated with Elon Musk’s Cybercab pitch: a vehicle owners could send into the world to earn money. Coogan recounted Marques Brownlee’s response to an earlier Cybercab event. Brownlee was impressed by the direction of self-driving, but skeptical that Tesla could meet the proposed timeline. He said he would shave his head if it did.
Tesla’s public description of Cybercab emphasizes the passenger experience: theater, music, and gaming applications on a 22-inch touchscreen, with future Starlink hardware planned. An interior tour shown by the hosts featured upward-opening doors and a screen-centered cabin. One viewer reaction was that the vehicle could replace Uber trips entirely. But Coogan’s framing kept returning to the operational conditions behind that appeal.
Musk’s cost claims are central. In a quote shared on X, he said Cybercab would cost roughly 20 cents per mile to operate, or 30 to 40 cents including taxes and other costs. He compared that figure with an average city-bus cost of about $1 per mile, excluding the subsidized ticket price, and described the result as “individualized mass transit.”
Coogan was bullish on Tesla’s current full-self-driving capability after recently testing a Model S and ordering a Model Y L. A colleague, Brandon, had said roughly 95% of his miles in the car were driven using full self-driving. Coogan’s view was that a purpose-built cab may avoid some difficult consumer-vehicle situations, including garages, driveways, parking lots, valets, and certain drop-off points.
Yet those remaining situations illustrate why the Cybercab promise is larger than hands-free driving. Coogan pointed to Tesla’s Smart Summon feature, which can bring a car out of a parking space to its owner, though he called it “a little janky.” The desired inverse feature, which he called “Banish,” would let a passenger get out at the curb and send the car away to find or pay for parking, then return later. That is the autonomy people actually want: not only driving the route, but handling the logistical residue of the trip.
Coogan argued that Cybercab makes more commercial sense than Tesla’s long-promised Roadster. A $250,000 performance vehicle may appeal to enthusiasts, but it is not a system for everyday transportation. A robotaxi, if it reaches the promised price and operating economics, addresses a substantially larger market. It also carries a substantially higher burden of execution.





