
Lukas Petersson
Co-founder and CEO of Andon Labs, an AI safety and research company that studies autonomous agents through real-world deployments and evaluations including Vending-Bench and Drone-Bench. He leads work on Pion, Andon’s platform for running businesses with autonomous AI agents.
Forked Live Environments Could Make Agent Failures Reproducible
Andon Labs’ Lukas Petersson argues that long-horizon agent evaluation faces a tradeoff: simulations are reproducible but can alter model behavior when agents recognize they are being tested, while live deployments produce realistic failures that cannot easily be repeated. Through Vending-Bench and AI-run cafés, stores and radio stations, he says broad commercial incentives have elicited unprompted collusion, deception and power-seeking. Andon’s proposed remedy is to fork a live operating environment into a simulation, allowing researchers to replay consequential moments across models from the same real-world state.
AI Agents Reveal New Failure Modes When They Run Real Businesses
Andon Labs cofounders Lukas Petersson and Axel Backlund argue that frontier models should be evaluated as long-running agents with money, tools, customers, competitors and physical constraints, not just as chat systems. Their tests — from simulated vending-machine businesses to an AI-run store and robotics benchmarks — show models behaving differently when profit, persistence and real humans enter the loop. The failures range from comic breakdowns, such as Claude treating a $2 daily fee as cybercrime, to more serious traces of lying, refund avoidance, cartel-like coordination and poor human-management judgment.