Orply.

A Small Research Lab Can Turn AI Experiments Into Products

Dan ShipperLenny's PodcastThursday, September 24, 20268 min read

At Lenny and Friends Summit, Every co-founder and CEO Dan Shipper argued that product teams should not be expected to explore fast-changing AI capabilities while also delivering a roadmap. He proposed a small research lab to test new ideas in real work, discard most of them and pass the few that prove useful to the product team for further development. The aim, he said, is to make room for experimentation without making customers wait while the whole organization chases each new model.

A product team cannot explore and execute as if they were the same job

When a frontier model arrives, product leaders face a choice that can feel absurd: keep executing the roadmap, ask customers what they want, or remake a business app as a 3D multiplayer strategy game. At Every co-founder and CEO Dan Shipper’s telling, none of those is a reliable answer on its own. Customers may not yet know what the new capabilities make possible; meanwhile, abandoning existing work to chase every new model is its own risk.

Shipper opened with examples of what recent models could do. He said Fable 5.1 and Astra 6 had launched the previous week. Using a single prompt, he had generated a historically accurate 3D simulation of the Battle of Waterloo; it took about four hours to produce. A colleague made a simulation with a thousand agents by giving a model a scientific paper about how agents interact, then asking it to visualize the result. Shipper also said Astra was capable of doing some of the company’s video editing in Premiere, and that it had made much of the talk’s animations.

The demonstrations make the management problem vivid: new capabilities can make an idea feel compelling before there is evidence that it solves a useful problem. Shipper’s first rule is not to make major life decisions within 30 days of a meditation retreat, a psychedelic experience, or a first encounter with a frontier model. The joke points to a real tension. Product leaders have to respond to what the technology makes possible without confusing novelty with durable value.

The deeper conflict is between two kinds of work. Exploring the frontier means trying many approaches, making demos, and expecting to discard much of what gets built. Executing a product roadmap means narrowing choices, saying no, and improving something customers already rely on. Shipper argues that asking one team to do both at once pulls it in opposite directions.

His proposal is to add research-lab practices to a product organization: protect space for exploration, then give the product team responsibility for turning the few promising results into coherent, scalable offerings. The aim is not to replace the product team with a lab. It is to let some people investigate what might come next without making everyone else abandon what they are already responsible for delivering.

Treat the team’s early adopters as a resource, not a mandate

Shipper’s model starts with a distinction already present inside many organizations: some people spend more time experimenting with new tools than others. He calls these people early adopters. They may be trying new models on weekends or using them for personal projects, and so may have an informed sense of what a product could become.

But their curiosity can also become a distraction. The management question is how to use what they learn without pulling the whole organization into exploration. Shipper’s answer is to separate responsibilities: a lab explores new capabilities, while the product team improves and scales what works.

A lab need not be a large, separate department. Shipper says AI can make a “labs team of one” viable: one person, perhaps assigned to investigate a new model when it appears, can try things and report back. He contrasts the lab’s appetite for experimentation with the product team’s obligation to customers. A lab should expect to throw away roughly 90% of what it makes; the product team should expect to adopt about 10% of what the lab tries.

90%
of lab work Shipper says teams should expect to discard

The distinction sets different expectations for the two groups. Exploration is allowed to produce failures; the product team remains responsible for delivering something coherent to customers. Shipper points to Anthropic Labs as an example of the structure. He attributes Claude Code, MCPs, skills, and Claude Design to a small group experimenting inside Anthropic, with further investment going to the ideas that worked. He says many experiments never became visible products. The public successes, in this account, sit on top of a larger body of work that was tried and discarded.

Keep the lab small, fast, and close to real work

For lab teams, Shipper favors one or two people rather than the conventional “two-pizza team” of roughly eight to ten. In his view, AI tools let a small group get far enough that the coordination costs of a larger team can outweigh its benefits.

He describes a useful pairing as “pirates and architects.” The pirate generates many rough attempts, looking for something valuable without worrying too much about polish. The architect takes a messy prototype and shapes it into something more coherent and extensible. Shipper casts himself as the pirate: someone willing to make a lot of things and throw most of them away. The pairing gives the lab both speed and the ability to turn a promising improvisation into a system.

The lab also needs a short feedback loop. Shipper’s preference is to build for oneself, because the builder can immediately tell whether a tool helps with real work or is merely novel. If that is not possible, he recommends working with a small number of early-adopter customers. Experiments should be tested against actual tasks, not just judged by whether a demo looks impressive.

That testing can mean building several versions of the same idea in parallel. To a product team, competing approaches may look incoherent; to a lab, they help map a frontier whose possibilities are not yet clear. Different people may find different ways to address the same problem, and comparing them helps the organization distinguish a useful capability from a striking novelty.

Shipper also argues that even experiments that do not enter the product can produce value. At Every, he says, the company turns some of its experiments into external content describing what it tried and what did or did not work. A lab can also use prototypes to bring early adopters closer to the company. At minimum, it should share what it has learned with the product team, so product decisions can reflect new capabilities without requiring the whole team to explore them directly.

Make adoption a sequence of tests, not a handoff by enthusiasm

Shipper describes a research pipeline that moves ideas through stages: lab-only experiments, internal use, early customers, and readiness to scale. Most experiments stop in the lab. The survivors are tested in real work, first by colleagues and then, if they prove useful, by early customers. The product team can then decide how an idea fits the roadmap and what it would take to release it more broadly.

His example is KateBench, an effort to help Every’s editor-in-chief, Kate Lee, with copy editing. Shipper says Lee has strong editorial taste, but as the company grew to around 30 people she could not personally do all the copy edits. Hiring someone with comparable judgment and training them would be difficult. He had therefore been experimenting for years with ways to use AI to extend her impact.

One experiment used Lee’s past edits to prompt Fable to copy-edit a new piece. Shipper would send her a draft and ask her to accept or reject the suggestions, producing more feedback. He says the idea only recently became useful enough for Lee to use herself. In an internal Slack exchange shown during the talk, Lee asks the company’s Every agent for a “Kate pass” on a draft; the agent returns suggested changes for her to accept or reject.

The next step was to make the rough tool more measurable and maintainable. Shipper says he brought in an architect, Yannick, to turn the prototype into a system. A dashboard tracked suggestions accepted or flagged, remaining manual work, and decisions made across documents. Shipper reported that Lee did 12% less work on those kinds of edits than in the preceding month.

12%
less work Shipper said Lee did on these edits than the month before

For Shipper, internal use is an important filter. If colleagues begin using a tool and return to it, that is a stronger signal than the excitement of a first demonstration. From there, a team can decide whether to expose it to early customers. Even at that stage, he says, not every idea should proceed.

The criteria for moving an idea forward should be explicit. Shipper’s questions include whether people use it and come back, whether it is “10x better” than what already exists, and whether it is affordable to serve at scale. He stresses the time dimension: what feels new may not feel better a month later. Usage over time helps separate durable value from the initial appeal of novelty.

Regular reviews make the pipeline visible to the product team. Shipper says Every tracks experiments in Notion and discusses them in weekly all-hands meetings. The point is to let colleagues responsible for the roadmap see what is moving forward and consider its implications without having to conduct all the experiments themselves. Clear decision criteria and recurring review keep the lab and product organization informed by one another.

A successful experiment can become the next product

The broader ambition is to let a company build a possible next version of its product while continuing to scale the current one. Shipper says the idea of disrupting one’s own product is not new; what has changed, in his account, is how quickly AI makes it possible to experiment and find a viable direction.

He points to Codex at OpenAI. According to Shipper, a small team worked outside the main app while other teams also explored the future of coding, including different product forms such as an IDE or a command-line interface. The Codex desktop app launched in February 2026 and, Shipper says, grew quickly enough that it was merged into ChatGPT and became a foundation for the app. He describes this as a small experimental team’s work scaling into a much larger product.

The measure of whether the lab model is working is not simply how many prototypes it produces. It is whether the organization can welcome a new model release as an opportunity to investigate, rather than dread it as another disruption to the roadmap. Exploration remains uncertain and most experiments will fail. The structure is meant to make that uncertainty manageable—and to give the few ideas that prove useful a path into the product.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free