Microsoft Bets Enterprise Agents Will Run Through the Cloud
John Coogan
Jordi Hays
Eric Glyman
Martin Scorsese
Satya Nadella
Steven BathicheTBPNWednesday, June 3, 202614 min readJohn Coogan reads Microsoft Build 2026 as a sign that Microsoft is trying to make the cloud, not the phone, the center of enterprise AI agents. On Diet TBPN, he argues that Project Solara, Scout, OpenClaw support and Microsoft’s own models point to a platform strategy built around Azure, Microsoft 365 data, security boundaries and cost-efficient deployment rather than frontier-model supremacy. The open question, he says, is whether agent hardware and workflows can win adoption outside environments where companies can mandate them.

Microsoft’s agent strategy treats the cloud as the hub
John Coogan read Microsoft Build as a sign that Microsoft is pushing deeper into its own AI stack, not only packaging AI through existing software. The company is “in the foundation model game,” he said, pointing to MAI-1, Code 1 Flash, and MAI Thinking 1 as Microsoft’s first coding and reasoning models. The pitch around those models was not primarily that Microsoft had overtaken the frontier labs. Coogan said speakers emphasized efficiency, cost per token, and return on investment: in an “ROI race,” the models have to be cheap enough to use.
The more consequential shift was Microsoft’s attempt to make agents native to its platform. Scout, Microsoft’s first proactive AI agent for Copilot, is built around access to the data that sits inside the Microsoft work stack: Teams, Outlook, OneDrive, SharePoint, chats, email, calendars, contacts, and enterprise files. Coogan’s read was simple: this is good news if a company is already “all in on the Microsoft ecosystem.”
That ecosystem framing also explains Project Solara, Microsoft’s Android-based operating system concept for devices designed to run agents instead of apps. Microsoft showed two Solara device concepts: a stationary desk device built on MediaTek silicon and a portable badge-like device built with Qualcomm. The desk device was presented as a companion for Windows PCs. It used Hello for Business so that a user could walk up to it, authenticate, and access agents. Microsoft’s stage language described the interaction model as “think, plan, delegate.”
| Project Solara concept | What Microsoft showed | Stated or implied role |
|---|---|---|
| Desk device | Stationary device built on MediaTek silicon, shown beside a Surface laptop | A Windows PC companion and secure desk-side interface for agents |
| Badge device | Portable Qualcomm-based device with a screen, user identity, fingerprint unlock, and agent tasks | A lightweight on-the-go access point for enterprise agents |
Satya Nadella introduced the broader Windows AI push by saying Microsoft was expanding the scope of Windows ML and Windows AI to tap into available compute power. The slides shown on stage described Windows AI APIs expanding to more PCs across CPU, GPU, and NPU, with on-device small language models, video super resolution, and speech recognition. Microsoft also announced Aion 1.0 Instruct and Aion 1.0 Plan as new models that would run “on Windows inbox.”
But Solara was the part Coogan treated as the real platform question. The device concept itself looked unresolved. The portable form factor, shown with the name Steven Bathiche on the screen, resembled a smart badge or key card more than a phone. The demo showed a user unlocking it with a fingerprint and seeing a task: “Gather content for your social media post for Friday,” followed by subtasks such as reviewing copy options and finding related files.
Jordi Hays asked the obvious question: why would anyone want this instead of an application on a phone? Coogan did not pretend to have a definitive answer. “I’m not sure,” he said. The most plausible case he offered was a “dumb phone” style inversion: a device that does not invite consumption but does let someone delegate work. You would not scroll TikTok on it, but you might trigger tasks to be completed later on a desktop.
The stronger argument came through Ben Thompson’s interpretation, as relayed by Coogan. Thompson described Project Solara as “vaporware at this point,” though Microsoft had shown real devices and named Qualcomm and MediaTek as chip partners. Still, he found the concept compelling because it changes the role of the device. Wearables are usually limited by their interaction model: they are useful only when the human is actively using them, and that interaction is often annoying or inefficient. Solara’s promise is the opposite: a brief interaction, followed by cloud-based agent work happening in the background.
Coogan summarized Thompson’s larger point this way: Microsoft has an interest in making the cloud the platform because it does not control the iPhone. But that self-interest does not make the model wrong. For agents, the cloud may be a better hub than the phone. Agents work across apps, devices, and large stores of enterprise context; phones are comparatively locked down. In an enterprise environment, the relevant context and compute are already in the cloud, and Solara is aimed at enterprise rather than consumers.
That makes the badge concept less absurd. A company could require employees to carry a secure badge that acts as an on-ramp to enterprise agents running in Azure and Microsoft 365. Coogan tied that to Thompson’s thesis that “thin is in” in the age of AI. If the scarce compute lives in the data center, then the client device can be minimal. The badge’s job is not to be a full computer; it is to be an interface.
Coogan compared the Solara badge’s form factor to Rabbit R1, noting that some people took Microsoft’s demo as vindication for Rabbit’s thesis: offload compute to the cloud, keep the hardware minimal, and let the device trigger useful work elsewhere. He remained skeptical of the adoption problem. People like phones. They like having a full-screen, general-purpose device that can do everything from work tasks to movies. Solara’s enterprise angle may solve that problem only where companies can mandate behavior.
OpenClaw fits Microsoft if Microsoft controls the ground it runs on
Microsoft’s embrace of OpenClaw was presented as strategically consistent with the Solara thesis. Coogan read from Alex Heath’s post, which said Microsoft had “buried the news” that Nadella was fully embracing OpenClaw. Scout, when released more broadly in the summer, would be powered by OpenClaw, and Microsoft would contribute its security guardrails back to the open-source ecosystem.
The post shown on screen framed the headline simply: “Microsoft embraces OpenClaw.” Coogan’s interpretation was that OpenClaw’s rough edges become less dangerous when the agent is constrained inside Microsoft’s walled garden. He referred to prior problems around OpenClaw and Meta AI, including Instagram account theft, as examples of what can happen when powerful agent tools operate without enough constraint. In Microsoft’s version, the agent can pull from spreadsheets, PowerPoints, databases, email, and other work artifacts, but within the permissioning and security structure of Microsoft cloud accounts.
The walled garden, in that reading, is not just a lock-in mechanism. It becomes a safety boundary. Coogan put it bluntly: if you are all in on one walled garden, “the walls are actually somewhat safe.”
Heath’s analysis, as quoted by Coogan, was that Microsoft is accepting an open framework because it believes it can control the environment in which that framework operates. “You only welcome a growing open framework onto your turf when you’re confident you can control the ground it stands on,” Coogan read. In this case, Heath’s read was that Microsoft is doing what it has historically done well: acting like a platform company rather than trying to own too much of the stack.
Apple was the contrast. Hays noted that Apple had been one of the primary beneficiaries of the OpenClaw boom because Mac Minis had sold out in places. Coogan agreed, but said Apple was not likely to embrace OpenClaw the way Microsoft had. Hays said Apple lacks the same enterprise motion; Coogan added security and privacy as further constraints.
WWDC, in Coogan’s view, might bring an OpenClaw competitor or a significant improvement to Apple Intelligence. His own expectations were more modest than an OpenClaw-style system that “worms its way into every other app.” He expected something closer to a Siri that finally works: basic shortcut integrations, near-frontier-level answers, and perhaps Gemini running under the hood. That would be useful, but different from Microsoft’s agent platform strategy.
Alex Heath’s other observations from Build added caution around the scale of Microsoft’s announcements. Coogan said Heath believed the Copilot “super app” was not ready; the new autopilot tab with Scout had been shown off, but not on stage, suggesting the broader rollout was delayed. Heath also argued that Microsoft was still not at the frontier in models. Coogan said Microsoft did not make a big deal out of benchmark dominance and appeared to compare some models to older Anthropic and OpenAI systems. The story was more about being competitive at attractive cost-performance levels than about having the best model in the world.
An unnamed participant argued that the more interesting comparison was not with OpenAI or Anthropic, but with Meta. MAI Thinking 1 was described as “quite competitive with Meta,” which stood out because Microsoft’s MAI team has not had the same public hype cycle as Meta’s AI efforts. Coogan contrasted Microsoft’s quieter presence with Meta’s aura-building around talent: Nat Friedman, Daniel Gross, Alex Wang, and high-profile recruiting created a narrative around Meta’s labs. Microsoft, by contrast, was producing “pretty solid models” without as much industry theater.
Microsoft is selling clean enterprise AI as much as model capability
The model discussion turned quickly from capability to liability. Coogan said Microsoft repeatedly emphasized a “very clean pre-training data set,” positioning its models as safer for enterprises worried about downstream legal or reputational risk. If a company does not want New York Times material or Harry Potter books in the training set, Coogan said Microsoft’s pitch is that those materials are not there because the company has done the sanitization work.
Microsoft also stressed that it did not distill from another lab’s model. Coogan treated that as a notable claim because accusations around distillation have circulated widely. He referred to the Elon Musk OpenAI lawsuit and said Musk had mentioned that he, or people at xAI, might have done some distillation on Anthropic or OpenAI models at one point. In that context, Microsoft’s “no distillation” position was presented as part of its enterprise trust package.
The product architecture for enterprise customization also mattered. Coogan said Microsoft launched features that let companies fine-tune models. He distinguished Microsoft’s approach from Amazon’s: Amazon offers mid-training, providing a checkpoint of the pre-trained model so customers can add business-relevant data and continue training from there. Microsoft, as Coogan described it, is offering more of a reinforcement-learning post-training step.
The common enterprise pitch is straightforward: Microsoft provides a baseline model with useful capabilities, optimized for Azure, with predictable price-performance and known deployment characteristics. Customers then adjust it for their own workflows and run it on the same hardware. They know what it will cost per token, they know where it runs, and they can make it answer their specific business questions more effectively.
Coogan did not claim adoption was assured. “At least that is the pitch, that is the hope,” he said. But he noted that Microsoft has a strong go-to-market and enterprise sales organization, which means the market will probably learn soon whether these deployments land.
That enterprise lens also shaped the jokes about “agentic commerce.” Coogan imagined a future in which agentic AI produces a $500 million bill, over a wire limit, requiring a transfer of “tens of thousands of Bitcoin” with low fees. Hays deadpanned, “It happens.” The joke sat on top of a real operational question the source kept returning to: once agents act on behalf of users and businesses, the surrounding systems for payment, control, and auditability matter.
Consumer hardware still has to beat the phone
The exchange about Chinese consumer brands sharpened a broader adoption problem: recognition and utility are not the same as admiration. Joe Weisenthal’s post, shown on screen, said Chinese running shoe brands were appearing more often in the subreddit r/runningshoegeeks: one out of 18 posts mentioned a major Chinese running shoe brand, up from one out of 40 in the prior quarter, with Li-Ning the most prominent.
Coogan’s point was that the brand itself, not merely Chinese manufacturing, was becoming more visible. Hays said Chinese companies have acquired a number of brands, including Arc’teryx, and that the next step is building meaningful consumer brands of their own. When Coogan asked for the Chinese consumer brand Americans like most, Hays answered DJI.
Coogan agreed. DJI has brand admiration in the U.S. in a way many Chinese marketplace brands do not. He described it as aspirational, closer to GoPro than to a generic Amazon product. He acknowledged that DJI is controversial because some people view it through an industrial-capacity or geopolitical lens. But for a consumer whose wedding video used a DJI drone, the brand association is positive and practical.
Temu and Shein came up as counterexamples. Hays suggested them as recognizable Chinese brands; Coogan granted their popularity but questioned whether they were admired in the same way. Temu, he said, feels more like a Walmart brand than a Nike or Apple-style brand.
That distinction helps explain the skepticism around Solara without turning the brand discussion into a direct proof point. A separate agent badge needs more than awareness. It needs either admiration, compulsion, or a use case the phone cannot absorb. Coogan’s earlier Solara analysis kept returning to that same tension: Microsoft may have a cleaner path inside enterprises, where a badge can be mandated, secured, and integrated with the work graph, than in consumer life, where the phone remains the default device.
Specialized experiences work when narrowness is the value
The Call of Duty demo offered a separate product-design test: when is novelty the point, and when is familiarity the product? Coogan described the system as generative level design, while clarifying that it was not GenAI in the current transformer sense. The demo showed a map assembled from “slabs,” modular square sections that can be nested and randomized at runtime. A speaker in the clip said the system had the content to support “upwards of 900” permutations.
The argument for the system was replayability. In a multiplayer game, players initially feel discovery when learning a new map, but that fades after five or ten plays. The clip claimed that with the randomized slab system, that discovery “never fades” for the map called Kill Block.
Hays joked that it posed an existential risk to TBPN because they might play it in the Ultra Dome and forget to go live. Coogan, however, was more conservative as a player. He said he keeps coming back to Rust, comparing it to a vacation destination full of good memories. He does not necessarily want the map remixed. Hays called him a Luddite.
The same specialized-hardware thread appeared in a brief discussion of an Anduril and NASCAR limited-edition VR racing simulator rig priced at $14,000. Coogan said Palmer Luckey was “back in the consumer entertainment VR industry” with the product, though not by making the headset itself. Hays said he had just put down a deposit on his own simulator three days earlier, but assumed the Anduril-NASCAR rig would be specific to NASCAR simulation.
The narrower point is useful next to Solara: not every dedicated device or constrained interface is a mistake. A randomized map, a racing cockpit, or an agent badge has to justify itself by doing something better than the general-purpose default.
Agentic traffic puts pressure on controlled workflows
Hays brought in a claim from Cloudflare’s Matthew Prince: bots had passed human traffic online for the first time in internet history. Hays said Cloudflare put bots at 57.5% of traffic, and Prince had expected that milestone later, first by the end of 2027 and then early 2027, before agentic traffic accelerated faster than he predicted.
Coogan’s explanation was that every agent task can fan out into many page visits. When a user “fires something off,” it may become hundreds of pages, visible in reasoning traces and tool calls. The source did not present a complete theory of web governance, but the implication for professional software was clear enough: if agents become routine, they do not merely replace clicks. They multiply background actions across the web, cloud services, databases, and internal systems.
That made the Ramp Stack demo a natural example of agentic work under controls. Coogan introduced a demo from Eric Glyman for Stack, Ramp’s AI operating system for accounting firms. Glyman’s framing was that every accounting firm is asking how to build AI into its process. One-off experiments, such as prompts for variance analysis or vibe-coded reconciliation tools, may work sometimes, but they break when security, auditability, and accuracy are non-negotiable.
The Stack interface shown on screen made that claim concrete. It included a prompt — “Good morning, Tina. Help me update the prepaid insurance schedule this month” — and a close workflow showing 38 completed tasks, 10 in progress, and none not started. Glyman said Stack can handle real work end to end, learn a firm’s way of working for each client, and always give the human final say before anything gets posted. Every action is recorded and auditable.
| Stack claim | What was shown or said |
|---|---|
| Security | Glyman framed Stack as one secure place to orchestrate AI coworkers for accounting. |
| Auditability | Every action is fully recorded and auditable. |
| Human approval | Stack gives the user the final say before anything gets posted. |
| Workflow specificity | The product can be taught a firm’s ways of working for every client. |
Coogan said this has been Ramp’s vision since its launch period: automate back-office finance work, but with the controls businesses require. Hays called it “god mode for the back office,” while Coogan called it “god mode for accountants.” The exaggeration sat on top of a concrete product claim: accounting AI cannot be a clever chat prompt if it is touching books, journals, and client work. It has to be orchestrated, permissioned, and auditable.
Scorsese gives image generation a workflow argument
Black Forest Labs showed Martin Scorsese using FLUX.2 in what the video called a “working storyboarding session.” The product text described FLUX.2 as a production-grade AI image generation and editing model with 4MP photorealistic output and multi-reference control.
In the clip, Scorsese described the image he wanted: a place that did not feel modern; a town, not a village or city; almost medieval; narrow cobblestoned streets; a main road twisting through the town; and a camera placed higher, looking down. He compared the process to Cecil B. DeMille having production designers create oil paintings.
This is that, in a sense. Conveys a cinematic, a cinematic intelligence.
Coogan said “cinematic intelligence” was a strong tagline for Black Forest Labs. He acknowledged that AI in filmmaking is deeply controversial, but argued that putting Scorsese in the workflow changes the conversation. The question becomes less whether AI will make the next Martin Scorsese movie by itself and more whether it can be useful as a tool inside a filmmaker’s process.
His answer was cautious but open. It probably will not make the next Scorsese movie this year, he said. But could Scorsese use it while thinking about what to work on next? “Sure.” That placed the FLUX demo alongside the enterprise examples as a workflow argument: AI is easier to evaluate when it is constrained by context, taste, permissions, and human approval.

