Agent-Ready Interfaces Can Override a Decade of Brand Loyalty
Liad Yosef, co-creator and maintainer of the MCP Apps specification, argues that online services will increasingly compete on how readily agents can discover, authenticate with, and operate them—not just on the quality of their human-facing dashboards. He says companies should test agents completing real tasks rather than rely on published standards such as `llms.txt`, which Ora found agents often ignore. In Yosef’s model, services become “nearly headless”: agents handle routine operations directly, while providers return an interactive interface in chat when human judgment is required.

A better agent interface can outweigh a decade of product loyalty
Liad Yosef puts the commercial consequence of the agentic web plainly: products will increasingly compete on whether an agent can operate them, not only on whether a person likes using their interface.
He described asking Claude Code to choose an analytics service for his team. The team’s preference was Mixpanel, which Yosef said they had used for a decade. Claude Code recommended PostHog instead. Its reason, according to Yosef, was not a better human-facing product experience but better MCP support and APIs for the agent to integrate against. The team accepted the recommendation and switched.
Mixpanel spent a decade perfecting their UX and developer experience, and we just left it just because Claude Code preferred PostHog.
The anecdote is Yosef’s case against assuming that established brand preference transfers to an agent-mediated market. An agent choosing a service to complete a task does not have the accumulated familiarity a person has with a dashboard, nor the patience to accommodate a product whose usable interface begins only after opening a browser. Yosef argued that if an agent has to spin up a browser to use a product, it will treat that as friction and may select a more directly operable alternative next time.
That makes “agent ready” a broader requirement than publishing an API or appearing in search. A service must be usable through several paths: its own agent, a customer’s general-purpose assistant such as ChatGPT or OpenClaw, browser-mediated automation, and conventional human browsing. The website remains one interface, but it no longer monopolizes access to the business.
Yosef’s readiness framework turns that proposition into an operating checklist. Discovery matters, but it is only the first condition. Once an agent finds a service, it has to identify what the service is, authenticate successfully, perform the requested work reliably, and return a usable human interface when a person needs to review or decide. The displayed lifecycle begins with intent, then moves through discovery, identity, authentication and access, integration, and user experience.
| Readiness dimension | Weight | Question Yosef poses |
|---|---|---|
| Discovery | 15 points | Can agents find you? |
| Identity | 20 points | Do agents understand you? |
| Authentication and access | 30 points | Can agents authenticate? |
| Integration | 20 points | Can agents do the work? |
| User experience | 15 points | Can humans use you when needed? |
Auth & Access receives the largest allocation, at 30 points. Yosef’s accompanying account makes authentication a necessary part of the path from discovery to action: an agent must be able to authenticate to a service and, in his formulation, eventually pay it headlessly. Integration is a separate question: whether the agent can actually complete the requested work.
But Yosef’s central warning is that a checklist can become compliance theatre if it measures what humans assume agents need rather than what agents actually use. His team’s work on llms.txt is the evidence for that warning.
Published standards cannot substitute for watching agents work
Liad Yosef presented an unexpected finding from Ora’s research: the signals companies publish for agents may not be the signals agents actually use.
His team scanned 12,982 sites and found that 5,462 of them—42%—published an llms.txt file. Yosef called it a de facto standard for signaling how agents should work with a site.
Yet the observed agent behavior did not make llms.txt the natural point of entry. In the displayed data, agents used documentation pages 86% of the time and homepages 84% of the time when those resources were present. llms.txt was used 40% of the time, while .well-known/* resources were used 34% of the time. Yosef said the agents that did reach llms.txt generally found it because the documentation told them it existed.
| Resource | Usage rate when present |
|---|---|
| Documentation pages | 86% |
| Homepage | 84% |
| llms.txt | 40% |
| .well-known/* | 34% |
For Yosef, the gap between adoption and use undermines a compliance-oriented approach to agent readiness. A team can publish llms.txt, auth.md, pricing information, and other standardized resources, but still fail if agents do not look for them in the course of completing a real task. The relevant question is not whether a file exists. It is whether an agent can discover the service, understand it, authenticate, and complete the intended work through the route it actually chooses.
He also argues that fixed best practices decay quickly as models change. Yosef said that an MCP-server description that might have needed three paragraphs six months earlier can now work in three lines. He attributed to OpenAI the position that it no longer publishes tool best practices because guidance can become obsolete as models improve. In his view, people cannot permanently define what an agent needs from a service.
We need the agents to define what the agents need.
Ora, Yosef’s company, has built a benchmark around those kinds of questions. He showed Attio receiving a 67-out-of-100 readiness score; the displayed assessment credited strong identity and access management while identifying a lack of developer-resource discoverability. The product returns agent-generated feedback as well as a score, reflecting Yosef’s view that the relevant test is whether agents can accomplish tasks rather than whether a human-designed checklist has been satisfied.
Ora Journey is the behavioral counterpart to that benchmark. The tool runs a selected agent against a domain and a specified intent, then records the path it takes in real time. Yosef showed one attempt to create an account on Telnyx taking 17 steps, and another task asking agents to locate Attio’s integration setup and getting-started documentation.
The value, Yosef said, comes from comparing many runs rather than treating one harness as representative. He showed Claude Code, Eve—described as Vercel’s harness—and ChatGPT given the same website and intent. Their paths were visibly different: one was deeply branched, while others moved through combinations of pricing pages, support paths, homepages, developer pages, APIs, and web search. Ora has run these journeys tens of thousands of times, Yosef said, to learn what agents seek and where they encounter friction.
A readiness score can organize work across discovery, identity, access, integration, and human UI. It cannot settle whether a real agent will find an auth.md file, understand an API page, or choose a competitor with a more usable MCP surface. Those are behavioral questions. Yosef’s position is that they need behavioral evidence.
The product becomes nearly headless, not invisible
Liad Yosef does not propose that all interfaces disappear. His model splits a service between an agent-facing operational layer and a human-facing approval layer.
For routine autonomous work, he said, an assistant may need no UI at all. A request to organize email can be completed in the background, followed by a simple confirmation that the work is done. But a person booking a honeymoon hotel may want to compare options, inspect a map, select a particular room or seat, or make a final judgment. That is the “last mile” where a service still needs to present an interface.
MCP Apps, the specification Yosef co-created and maintains, is intended to support that split. It allows an MCP server to return an interactive UI component inside a chat rather than reducing the service to text or a database-like response. Yosef showed a conversational planning flow in which Booking.com lodging information and AllTrails recommendations appeared together in one chat context.
The arrangement gives a specialized provider a role that a general-purpose assistant cannot simply replace. In Yosef’s example, an assistant can understand the user’s calendar, preferences, and desired trip, but it does not itself negotiate hotel inventory or resolve a venue customer’s seat-change problem. The provider retains those capabilities and can return its own recognizable interface for the final decision.
I don’t want to use your agent. I want to use my agent.
That preference for a personal assistant is central to Yosef’s critique of the browser as the default agent interface. He argued that browser controls—filters, pagination, sorting, dashboards, and vendor-specific flows—were optimized over decades for human perception and navigation. A person planning an anniversary might have to move among a calendar, a marketplace, hotel sites, and booking interfaces, repeatedly translating the same intent into different systems.
The assistant, by contrast, can carry context across those services. Yosef’s example was an assistant that knows an anniversary is approaching and knows the user prefers nature-oriented hotels. It can retrieve the relevant hotel map, coordinate with the calendar, and surface relevant products without requiring each provider to build every possible cross-service integration. The user does not need most of Booking.com’s or Airbnb’s dashboard, he argued; they need the subset of capability necessary to express and execute the current intent.
Yosef does not dismiss browser agents as useless. He described Google’s WebMCP approach as a short-term bridge: instead of forcing an agent to infer a webpage from screenshots, a site can expose JavaScript tools that the agent can use while co-browsing. That can improve an assistant’s ability to find a deal or act on a page the person is already viewing.
His objection is to treating that model as the destination. A browser agent operating a Salesforce dashboard still depends on a dashboard designed for a person to navigate. In Yosef’s preferred model, the agent should work directly with the service, and the service should return a purpose-built UI only when human judgment is needed.
He calls this the “nearly headless” web. The underlying service interaction is headless: the agent uses operational interfaces rather than visiting a conventional site. But it is not completely headless because a human eye remains at the end of consequential workflows. Yosef describes the resulting design problem as the combination of agent experience—how an agent encounters and operates a product—and user-agent experience—how the person experiences the agent’s work and approval interface.
Yosef predicts that assistants will become the primary entry point to more online services, while qualifying that websites and browsers will not literally vanish. He cited people who connect Jira’s MCP server to an IDE and stop visiting Jira’s website, as well as his nine-year-old using ChatGPT rather than Google for searches. He also pointed to David Cramer’s argument that Sentry needs an API-first surface and may eventually see its web application become a minority of customer usage. In Yosef’s account, human interfaces do not cease to matter; they become one endpoint in a service model increasingly initiated and navigated by an assistant.
Agents need a discovery layer built for operational resources
Liad Yosef argues that behavioral testing on individual sites leaves a larger infrastructure problem: agents need a way to discover the resources through which they can use those sites.
He frames discovery as a recurring feature of technical shifts. The web had search, mobile had app stores, and social platforms had feeds. The corresponding resources for an agentic web include MCP servers, APIs, openapi.json files, documentation, and other operational surfaces. The question is not simply which company has the most recognizable homepage. One travel service might offer a stronger MCP server while another has more useful APIs; an agent needs a way to find and compare those capabilities.
Yosef considers conventional web search inadequate because it is based on human SEO and page rank, and is not flexible enough for agentic resources. Per-agent registries are fragmented and require each service to register repeatedly, he said. A central registry raises different problems: who curates it, what governance model applies, and how the registry decides which resources belong.
He pointed to emerging standards including ai-catalog.json, which he said standardizes how a website exposes itself to agents, and Agentic Resource Discovery, intended to standardize how a discovery layer or directory exposes resources to agents.
Ora.directory is Yosef’s proposed implementation. He said it indexes domains Ora has scanned, exposes ai-catalog.json files, and allows agents to query by capability or domain. In a displayed example for monday.com, the generated catalog identified an MCP server; another screen for Vercel listed agentic resources including an MCP server, llms.txt, documentation, and agent skills. Yosef described the directory itself as compliant with Agentic Resource Discovery, so an agent can query it directly.
The goal is not to replace every site with a single intermediary. It is to give agents an index of usable, scored operating surfaces—an answer to the problem that an agent cannot reliably discover a business merely by treating its human-facing website as the authoritative interface.
Accessibility signals serve both people and agents
Liad Yosef closes with a practical overlap between agent readiness and accessibility. Language models, he said, do not visually experience interfaces, buttons, and layouts in the way sighted users do. They depend on other signals—language and structure—to understand what a site is, what its controls do, and how to proceed.
In that respect, Yosef compared agents to people with visual disabilities. Improving a site’s accessibility for human users can therefore make it easier for agents to interpret and operate; work that helps agents can also improve human accessibility. The point is not that the two users are identical, but that both depend on a service exposing meaning beyond its visual presentation.
His broader conclusion is not to rebuild the web solely for agents. It is to make the existing web agent-accessible: able to be discovered, understood, authenticated against, and used operationally, while retaining a clear interface for the moments where people need to make the decision.


