Orply.

AI Product Differentiation Is Shifting From Interfaces to Intelligence

Sonya HuangSequoia CapitalTuesday, August 11, 20269 min read

Sequoia Capital partner Sonya Huang argues that companies should decide which AI capabilities to own down to the model weights and which to rent from frontier providers, rather than treating sovereignty as an all-or-nothing choice. Her test is whether cost, latency, domain-specific performance and proprietary data make intelligence central enough to control. As open-weight models approach frontier performance, Huang says application companies can use their data, evaluations and production feedback to build specialized systems that outperform general APIs in their domains.

The decision is not whether to own AI, but where

Sonya Huang defines sovereign AI as owning intelligence “down to the weights,” without external dependencies. But she does not treat it as an all-or-nothing proposition or as a mandate to abandon closed-model APIs. Opus and GPT remain useful, she says, for frontier-level APIs, desktop work, and coding agents. The strategic task is to decide which capabilities are sufficiently central to own and which are sensible to rent.

Huang’s test has four variables: cost, speed, performance, and the proprietary nature of the data involved. A company is not either 0% or 100% sovereign; it can use external models where their out-of-the-box capability is valuable while building and controlling the parts of its intelligence where economics or differentiation demand it.

Decision factorMore likely to ownMore likely to rent
CostAI is a high-cost line item in COGSAI cost is low
Speed and latencySpeed is a priorityLatency is not a P0
PerformanceA lower initial floor but higher domain-specific ceiling is availableStrong out-of-the-box performance is the priority
DataThe relevant data is highly proprietaryThe relevant data is less proprietary
Huang’s framework for deciding which AI capabilities to own and which to rent

Cost matters most for companies with low, zero, or negative margins, Huang says. The paradox is that a successful AI product can make its own economics harder: greater usage produces higher AI cost of goods sold. For companies with frequent model calls and thin margins, she frames ownership less as a technical preference than as a necessity.

Latency creates a separate reason to own. In coding and security, Huang says, a smaller distilled custom model can beat a larger general model when response time matters enough. Her coding example draws the distinction sharply. Coding agents are still largely rented because users want strong capability immediately and latency is not necessarily the top constraint. Tab autocomplete has a different interaction pattern: calls are frequent, speed matters intensely, and API costs accumulate. Huang says most tab-autocomplete models now run on sovereign intelligence.

The performance calculation has changed as open-weight models have improved. Huang says that companies once accepted a tradeoff: open weights offered control, while closed models offered the best performance. Now, in some domains, a company can tune an open model on its own data and exceed the performance of a closed API for that use case. She points to Kimi K3 and GLM 5.2 as especially capable open-weight starting points.

The final consideration is independence. Huang describes Anthropic and OpenAI as strong partners to the ecosystem, but says companies increasingly want “their own set of independent legs to stand on as well.” A closed provider can remain part of the stack without being the sole owner of a product’s most consequential capability.

Not your weights, not your product.

Sonya Huang · Source

Huang borrows the formulation from the crypto-era slogan “not your keys, not your coins.” Her version is not a claim that every model weight must sit inside every company. It is an argument that where intelligence is core to the product, a company should be able to control and custody the weights that produce it.

Product differentiation is moving into the intelligence layer

Sonya Huang argues that the contest between foundation-model labs and application companies is no longer only about the application layer: the interface, distribution, customer relationship, and product wrapper. It is increasingly a contest over the intelligence inside the product.

In Huang’s framing, labs approach that territory “tech-out,” extending from models toward the user-facing product. Application companies move “customer-back,” beginning with a customer problem and building toward the underlying intelligence. The question is no longer simply who controls the interface, but who can shape better intelligence for a particular workflow, customer base, and body of proprietary information.

“The product is the intelligence,” Huang says. That is why she describes application companies as “the newest neo labs.” She points to work from Harvey, Factory, Glean, OpenEvidence, Semgrep, and Ramp across evaluations and benchmarks, harness engineering, fine-tuning techniques, and other applied research. These companies, in her account, are producing technical work aimed at the specific problems their products encounter rather than merely packaging general-purpose models.

Huang situates that commercial shift within a broader preference for decentralized intelligence. Her alternative to a single opaque system powering more of the economy is a common foundation on which people and companies build specialized intelligence: systems shaped by their own data, industries, personalization, ways of working, and taste. Closed frontier models can remain useful within that world. What she rejects is outsourcing intelligence wholesale when it has become a core part of the product.

Ownership requires an offensive research function

Sonya Huang says the organization responsible for this work should not simply be a conventional AI platform team with another mandate added to its backlog. Many companies have historically organized AI in a hub-and-spoke structure: a central platform group supports multiple product or application teams. She advises against shoehorning that group into sovereign-AI work.

Owning intelligence is not, in her view, primarily a platform-services capability. It calls for people who can make decisions at the frontier, conduct research, and work offensively rather than only service internal teams. Huang’s recommendation is to begin with a small, distinct, de novo team—potentially a lab with its own mandate.

There is no single required background for its leader. Huang contrasts Harvey research leader Niko Grupen, whose experience includes research roles at Apple and Google Brain, with Ramp Labs leader Alex Shevchenko, whose background is more engineering-oriented, including roles at Microsoft and Ramp. The appropriate profile depends on the kind of work a company needs: more fundamental research, more applied research, or some combination.

The case for starting small is not merely theoretical. Huang cites Harvey, which she says has published substantial research with a team of seven. The point is not that seven is a universal target; it is that a focused group can make meaningful progress without recreating a frontier foundation-model lab.

Public visibility is part of the function, too. Huang calls “legibility” an underestimated strategic requirement. Companies may be doing meaningful work internally, but buyers cannot reward technical sophistication they cannot see. In a market where buyers are choosing an AI champion and hearing similar vendor pitches, they are trying to infer which company has the technical depth to carry them into AI adoption.

Huang invokes Winston Weinberg’s formulation that a CEO has two responsibilities: drive substantive results and control the narrative around what is being built. For an AI company, she says, those responsibilities are connected. Technical marketing, a separately branded research group, and carefully published research can establish that a company is doing more than making the same external API calls as its competitors.

The roadmap begins with strategy and evaluations

Sonya Huang puts strategy first: establish the capabilities the company intends to own and the capabilities it will rent. The next task is evaluations. She calls evals unglamorous and not particularly fun, but says the work done early determines how effectively a company can make every decision that follows.

Evaluations are the reference point for the rest of the work. They let a team assess whether a new router, harness, prompt, context system, training technique, or model choice is actually improving the system. They also support monitoring for performance drift in production.

Area of workWhat Huang says companies may doWhy it matters
StrategyDecide which AI capabilities to own and which to rentSets the scope of the technical and organizational investment
EvalsDefine measures of quality before major model workProvides the basis for judging later changes
Harness, routing, prompts, and contextExperiment with system design around existing modelsSome companies can achieve strong results without training a model
Post-trainingTune models for domain-specific gainsCan improve performance where the company has useful data and a clear target
Mid- or pre-trainingMove deeper into model development in rarer casesMay be necessary for some companies, but is not the default path Huang describes
Online learningTurn live customer interactions into a feedback loopLets the intelligence improve with use
The building blocks Huang identifies in an owned-intelligence roadmap

Huang does not present this as a uniform sequence every company must follow. Companies can find strong performance through out-of-the-box models, routing, harness engineering, prompts, and context. Others see gains from post-training. In rarer cases, she says, teams may need to move into mid-training or pre-training.

Her broader claim is that near-frontier open weights make the upper end of this work practical in a way it was not previously. Closed APIs offer a higher initial floor: a company can use a capable model, provide prompts and context, run evaluations, and build useful applications without collecting large training datasets or training a model itself. Huang characterizes that arrangement as a lower ceiling because the company cannot use its domain data and production interactions to improve the underlying intelligence.

Open weights require more operational and technical work, but Huang says they are more malleable. A company can start with an open base model close to frontier capability, then use post-training, prompt and harness engineering, and online learning to pursue stronger performance in its own domain.

You start with a baseline that’s already close to frontier, and then with a good enough technical roadmap, you can actually reach better than frontier performance by owning your stack.

Sonya Huang

An owned stack turns a model call into a feedback system

Sonya Huang divides an owned AI system into a production stack and a development stack. Production serves the user in the moment; development measures, trains, and improves the intelligence over time. The connection between them is central: customer interactions can become signals for improving future performance.

LayerCore components Huang identifiesOperating purpose
Production: harnessHarness logic, routing, promptsCoordinates how the system uses the model for a task
Production: tools and contextTools, vector databases, knowledge graphs, MCP connectorsMakes relevant information and external capabilities available during inference
Production: modelBase model, post-training, and in some cases pre-trainingProduces the user-facing intelligence
Development: measurementEvals and production monitoringMeasures quality and watches for performance drift
Development: improvementDomain-specific data and online learningUses trajectories, synthetic data, RL environments, and live interactions to improve the system
The production and development architecture Huang presents for an owned AI stack

The production stack is more than a model. Huang describes it as a harness on top of a model, with tools and context around it. In a closed-model environment, that can remain comparatively simple: use a model such as Opus or GPT, work with an accompanying harness, add custom prompts and context, build additional harness logic if needed, and run evaluations. A company can get far with that architecture without collecting large volumes of training data or training a model itself.

Owning intelligence makes the production system more configurable and more demanding. Huang describes the move away from a clean API call as opening “Pandora’s box.” A company must choose an open base model, determine whether and how to post-train it, select and configure a harness, define the system’s tool use, and construct the context available to the model.

Context is a material performance variable in this design. Huang identifies vector databases such as Turbopuffer, enterprise knowledge graphs such as Glean, and open-source connectors through MCP as different ways of making information available to a model. She also points to Engram’s work on encoding context in the weights themselves.

The development stack becomes correspondingly more important once a company owns the model layer. Teams need evaluations both to compare changes before deployment and to watch for drift in production. They need high-quality domain data for post-training, which Huang says can include expert trajectories, synthetic data, or reinforcement-learning environments. They also need online-learning systems that turn live customer interactions into feedback for future improvement.

A Harvey Research post shown during Huang’s remarks illustrates the visible research program she has in mind. The post announced Harvey Research, said the company had open-sourced Legal Agent Benchmark, and described it as the largest benchmark for long-horizon legal work, spanning 1,200 tasks across more than 24 practice areas. It also listed collaborations with Baseten, Trajectory, LangChain, Fireworks AI, Applied Compute, and Engram. The example joins the technical and commercial parts of Huang’s argument: an owned stack requires work on evaluations, data, inference, and agent design, while legibility requires making that work visible.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free