Orply.

Enterprises Are Building AI Strategies Around Model Independence

Alex AtallahAmjad MasadErik Torenberga16zSaturday, October 3, 202616 min read

OpenRouter’s Alex Atallah and Replit CEO Amjad Masad argue that enterprise AI is likely to depend less on one all-purpose model and more on a mix of specialized systems that companies can choose, combine and control. Atallah makes the case for keeping model options open; Masad says businesses also need independence from model, cloud and data providers. Both see narrower models as potentially cheaper and easier to govern, though the speakers leave unresolved how agents should coordinate safely and when specialization beats a general model.

Independence is becoming part of enterprise AI strategy

Alex Atallah describes OpenRouter’s purpose as helping companies build AI products without becoming locked into a single model or provider. That means more than making it easy to switch vendors. Companies, he argues, need to combine models trained in different ways—sometimes including models of their own—with their data and other services. Atallah calls this “neurodiversity”: using differences among models to build a product that can do more than a customer would get by using ChatGPT or Claude directly.

There is an economic case for keeping options open, too. Some businesses will not emerge unless AI becomes affordable enough for their use case. Atallah sees a marketplace as a way to create pressure on price and make it easier to discover which models work for which jobs. Before OpenRouter, he says, AI was effectively a one-player market. In a field where capabilities are hard to assess from a list of features, companies need to try models in real workflows to learn what they are good at. OpenRouter’s role, in his account, is partly to make that exploration possible and help companies stay near the “Pareto frontier” as the ecosystem grows.

That case for independence also shaped how Atallah describes OpenRouter’s acquisition by Stripe. He says the companies had stayed in contact for years through work streams and events, but an acquisition was not something OpenRouter had been considering before Stripe reached out. The discussions moved quickly. Atallah says Stripe was “founder friendly” and that it was important to him that OpenRouter could retain autonomy over its brand, roadmap, and product while moving faster and building a more serious go-to-market plan.

He also describes a shared purpose: building a neutral, trusted, developer-friendly platform that businesses can rely on, while making it easier for new companies to start and grow. Stripe and OpenRouter, he says, do not want everyone to become part of one giant company. They want reliable infrastructure, efficient markets, and streamlined workflows that can support both lifestyle businesses and venture-backed companies. Atallah sees payments and inference as services that may increasingly blend together for future businesses.

The scale of those ambitions came up when Amjad Masad discussed the potential markets that large technology companies see before them. Referring to a SpaceX TAM chart shown during his remarks, he pointed to its estimate of $26.5 trillion for AI, alongside $370 billion for space and $1.6 trillion for connectivity. He contrasted the total $28.5 trillion estimate with a world economy he described as roughly $100 trillion. Masad’s point was not that those projections were certain, but that companies with ambitions on that scale may be difficult to treat as neutral partners.

$26.5T
AI market estimate in the SpaceX TAM chart shown during Masad’s remarks

Masad says companies that work closely with foundation-model providers face the possibility that those providers will move into their customers’ businesses. He cites Figma in relation to Anthropic and Harvey in relation to OpenAI as examples of the concern. In his view, the model companies see a large share of the economy as a potential market, which can make partnership difficult.

Replit’s response, Masad says, is to become an “independence layer” for enterprises: an intermediary between a company and the models, but also between the company and cloud and data providers. The aim is to find the best token at the cheapest price while allowing customers to deploy across services such as AWS and Azure and use platforms such as Databricks and Snowflake. He argues that companies need more ways to gain independence across technology, not just AI.

Atallah says he was surprised by how willing enterprises were to try open-weight models. He expected brand trust to steer companies toward providers that other large organizations already used. Instead, he saw businesses exploring alternatives for cost and differentiation, and looking to build some of their own AI capability. AI, in his account, is no longer a technology project that can be marked complete after an initial rollout. Boards keep asking what comes next, and internal teams need a continuing strategy.

That makes evaluation part of the work. Atallah says he has been surprised there have not been more company-built benchmarks, though he expects that to change. Companies need ways to test models on their own tasks. He points to Masad’s work at Replit on cost per task and experiments with agents as examples of internal research that other companies may undertake. The aim is not simply to find a model that performs well in general, but to understand what works for a company’s own tasks and what it costs.

Masad compares this need for internal capability to the way companies came to need people who could build websites and software. He argues that every company will need an AI practice, and that the knowledge accumulated inside it—about use cases, model fit, and cost—will compound over time. That capability gives an organization a basis for making choices as models and prices change, rather than treating AI as a single product decision made once.

The enterprise challenge is not just selecting models. Masad says that making AI products do real work inside companies remains unresolved, in part because of data sovereignty and security. Consumer agents may be able to connect to personal accounts, but he says companies are more protective about enterprise data and the ways it might leak. Replit spent nearly a year, he says, working to make its product deployable on a customer’s own cloud, or on-premises. He says that was not what he expected to be doing when he believed cloud software would be the clear direction.

For Masad, the work of making AI useful inside an enterprise includes giving customers control over where the system runs and what services it uses. He says consumer use cases may look straightforward, while the industry still has substantial work to do before AI systems become useful and productive at work. Independence, in this account, is not only the ability to switch models. It also involves control over deployment, data, and the company’s own ability to understand its AI systems.

General agents can connect information but blur responsibility

The appeal of a general agent is its ability to join information across domains. Masad describes a Replit tool he first built as a CRM agent, which gradually expanded as he connected more of his information. With access to personal chat history, GitHub, Salesforce, and calendar data, it can connect details across otherwise separate domains—for example, linking a past meeting with a person to a current discussion between that person’s team and Replit’s sales team. Those connections, he says, can help him prepare for meetings and move deals forward.

Masad sees a real benefit in this kind of cross-domain context: a system can find relationships that would be hard to notice by looking at each source separately. But he also says there are cases where a focused agent is preferable, because a user may not want one system to do everything. The tension is between the synergies of bringing information together and the need to limit what a system can access or take responsibility for.

Atallah takes that tension further. The more work a general agent takes on, he argues, the more understanding its user gives up. If people become less stressed because an agent handles some area of work, he says, the agent does not itself take on the responsibility or stress that a person would otherwise carry. The user may delegate activity without gaining a clear account of who is accountable for what happens.

Atallah’s own general agent scans each day for things that need his attention and tries to decide what to do. Yet he finds it difficult to improve: after making changes, he tends to ignore its output again about a week later. The difficulty, as he describes it, is not only whether the agent can do the work. It is also whether he can trust its attention and judgment across many separate areas, and adjust how much he relies on it in each one.

His proposed alternative is a set of vertically focused agents, each responsible for a defined area and subject to quality checks, perhaps coordinated by a chief-of-staff agent. He illustrates the difference by comparing one capable chief of staff who drafts replies across an organization with ten equally capable chiefs of staff, each responsible for a separate part of someone’s life. The second arrangement, he says, makes it easier to decide where to rely on an agent and where to remain closely involved. Specialization is a way to give the user a more legible boundary around delegation.

Masad recognizes this as a return to specialization. Drawing on Adam Smith’s account of specialization and the division of labor, he argues that specialization can make machines more useful. But he cautions against treating specialization as an uncomplicated good for people: when human work becomes too narrow, workers may lose sight of its purpose and the results of their labor. Referring to the Marxist idea of alienation, he suggests that people can become detached from the broader product and the impact of their work. His tentative distinction is that humans should remain generalists while machines become more specialized.

The product question remains open. Atallah says he has not yet seen a specialized-agent system as elegant to use as a general assistant such as ChatGPT, Claude, or Muse. The design challenge is to make separate agents useful without forcing a person to manage a collection of awkward, disconnected tools. One pattern he had seen in discussions of Grok bot was to use separate bots with separate credentials—for instance, one for a bank account and one for a social account—that could communicate to complete a task without sharing credentials.

Masad says consumer agents and work agents may have different requirements. At work, employees’ access permissions limit how much context an agent can have, even if a CEO with administrative access can use a more general agent. A broadly capable agent may therefore be inappropriate for an employee who should not see every part of the company’s data. Specialization can fit those permission boundaries, but it also raises the question of how agents should exchange information when a task crosses them.

Agent collaboration needs protocols and checks that are still unsettled

Masad says there are not yet good protocols for agent-to-agent communication, and doubts that natural language alone is the right medium. Agents may need another protocol or a domain-specific language, alongside ways to preserve data isolation. His concern is that an agent might persuade another to provide information it should withhold.

He refers to a Hugging Face hack in which agents began helping one another, and says it seemed to him that the next generation of OpenAI models was trained for agent collaboration. But coordination creates its own security question: even if agents can work together, how can a system ensure that collaboration respects the boundaries set by users and organizations? A protocol would need to support useful communication without allowing one agent to obtain credentials or data that it should not have.

Atallah proposes a possible check between agents and between an agent and the infrastructure it uses. He says OpenRouter has a small internal prototype using a fast decision model to review tool calls or messages. In his example, the model looks at the original system prompt and a proposed tool call, then checks whether the action fits the agent’s instructions and additional guidelines.

The example is a red-team task in a sandbox. Agents are asked to test a product and should not access the internet. If they are told to stop immediately if they do, that instruction could affect how they attempt to break out of the sandbox. Atallah suggests that a separate model could check their actions against the restriction without disclosing it to the agents performing the test. The prototype is an idea they are trying internally, not a demonstrated or validated safeguard.

Atallah describes this as a possible use for a cheap, fast decision model: it could classify an action and give feedback when a tool call should be rejected. He says such a model might help bridge agents to one another and to infrastructure. Masad asks whether the approach amounts to policy enforcement, and Atallah says that it does.

Atallah also refers to an Nvidia open-agent-safety project, which he recalls as “Open Shell.” He presents it as an example of structural safeguards that companies might combine with model-based checks, while qualifying his recollection of the project’s name. Neither the internal prototype nor the recalled Nvidia example is offered as proof that the problem has been solved. The conversation leaves open how to make the checks reliable and how agents should communicate under them.

Narrow models may be cheaper, more controllable, and easier to maintain

Amjad Masad proposes that general models could train their own replacements for narrower jobs. He compares the idea to a just-in-time compiler: as a system runs, it identifies an opportunity to optimize and generates more efficient code. In the analogous AI setup, a general model or an observing agent recognizes that a recurring use case is limited enough to handle with a purpose-built model. Masad argues that such a model could be cheaper and, because it is less capable, less vulnerable to prompt injection and less able to cause harm. He presents this as a possibility, not an established result.

The proposal is not limited to classifiers. Masad says a replacement could generate unstructured text or make structured decisions. If the inputs and task are known in advance, he suggests, a team might take an existing model such as Qwen and train it to apply a particular policy. Companies may be using highly capable foundation models for work that does not require that level of capability—“nuking a butterfly,” as he puts it. The potential advantage, in his account, is not simply a lower bill: a narrower model may have fewer capabilities to misuse.

Atallah asks whether these models would handle open-ended text or structured decisions. The exchange leaves room for both. He initially understands the cost argument, while Masad emphasizes safety as well. Atallah acknowledges the point, but notes that frontier labs may produce very low-cost models that companies can switch to. The discussion does not resolve when training a task-specific model would be preferable to switching to a cheaper general model.

At Replit, Masad says he has trained small models for specific internal jobs. One estimates the cost of a prompt by producing a probability distribution over price buckets, such as $5–$10 and $10–$20. He describes this as a classifier trained using enumerated outputs and their log probabilities, a technique he has used for years, including to train a chess bot. Replit’s large store of internal data makes such task-specific models practical, he says.

Atallah calls this “less model debt.” Companies may hesitate to fine-tune a model for open-ended output because its capabilities can quickly feel out of date, requiring the work to be done again. A bespoke classifier has a narrower job: it does not need to keep up with new languages, programming languages, or the wider capabilities used to evaluate general models. That narrower scope may make it easier for a company to build and maintain without constantly worrying that it has fallen behind.

Masad likens the cycle to the adoption of dynamic programming languages. Python, JavaScript, and Ruby made it easier to build quickly, he says, but teams later confronted performance problems and bugs, then added tools such as types and just-in-time compilation. He expects a similar progression with general-purpose AI: companies will first use capable models for many tasks, then recognize that some applications are unnecessarily costly or risky and reach for narrower models. In his forecast, a company could upload a CSV and generate a specialized model for a single task.

That shift would also restore some of the predictability that software has historically offered. Masad says people may come to miss the period when computers did exactly what they were told. He points to DSPy as a sign of interest in more controllable model outputs, and argues that combinations of specialized models may accomplish more than expected. He acknowledges the flexibility of a promptable foundation model; the question is which tasks need that flexibility and which can be handled by a narrower system.

Greater intelligence does not settle whether models will deceive

The safety case for smaller models depends partly on whether greater intelligence makes systems more reliable or more dangerous. Alex Atallah says public evaluations are lacking, with many relevant assessments kept private. He sees arguments in both directions: smarter models may become better at alignment and coordination, but risks may also rise as models gain capability. He cites a claim by Noam Brown that agents have gotten better at coordinating as they have become smarter, while emphasizing that the implications remain uncertain.

Masad is skeptical that intelligence alone will make models less deceptive. He invokes the orthogonality thesis—the idea that intelligence and morality are separate—and says he does not think the idea applies straightforwardly to humans. In machines, he worries that reinforcement learning could make reward hacking and deception more effective. He also says that monitoring a model’s chain of thought can create pressure for it to lie in that chain of thought. He presents that as a concern based on what he says has been shown, not as a settled conclusion about how all models behave.

To assess deception, Masad argues, it may be necessary to run models for months on large tasks, rather than rely on brief tests. He prefers to discuss the specific problem of a model deceiving its user rather than use “alignment” as if it had a settled meaning. The term raises questions about whose values a system should follow, he says. Atallah frames the open question in practical terms: whether a sufficiently capable model will reliably stop deceiving or sandbagging during training, and whether anyone knows how to predict that behavior.

The speakers leave open whether increasing capability will improve or worsen the problem. Atallah says some arguments suggest alignment could improve as models get better, while Masad is concerned that models may become better at hiding reward hacking or deception. Neither claims to know what will happen. Their disagreement turns partly on how much confidence can be placed in evaluations, especially when a model might recognize that it is being evaluated.

If organizations could identify a model that was significantly less deceptive, Atallah suggests, some might choose frontier models even at much higher cost. The price would depend on the task. He says organizations might spend ten times more for a less deceptive model when conducting security research or code review, which he describes as especially high-risk applications.

Atallah also argues that structured decision tasks may offer a different risk profile. When outputs are constrained, machine-consumed, and not executable code, he says, the room for misbehavior is smaller. He expects enterprises to take more interest in those tasks, which he says are underrepresented in discussions of model use. This is a reason to examine whether a task needs open-ended generation at all, not a claim that structured outputs eliminate risk.

Model fusion has produced specific cost and performance claims

For Atallah, model diversity is not only about assigning different jobs to different systems. It can also mean combining model families in one workflow. He says model-fusion research had progressed slowly for years, but that work was accelerating. Code review had been an early example of using different model families to double-check results. OpenRouter and Cognition had also launched fusion tools, he says, with OpenRouter’s initial focus on deep research.

The rationale is that model labs train on different sources of data. A system that draws on more than one model may search a wider range of ideas instead of relying on a single provider. Atallah says OpenRouter’s initial result delivered frontier-level quality at half the cost. That is his description of a particular result, not a general claim about fusion systems.

Masad says Replit had published a result that day for a “deep sweep” through the Replit agent, combining several elements, including the harness—the surrounding system that runs the agent. He describes the result as frontier-level performance at 40 to 50 percent of the cost. The exact models used change over time.

Atallah asks which models the system used. Masad says the mix changes, then offers a tentative explanation of efficiencies within the OpenAI model family. He believes a feature may allow computation to be reused across different models or effort levels, but explicitly qualifies that he might be wrong and says he would need to double-check.

Atallah adds that cache-awareness is important when designing fusion systems, routers, and escalation mechanisms. The speakers’ examples are specific results and design observations; they do not establish that every fusion approach will achieve comparable performance or savings.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free