
Ronak Malde
Co-founder and CEO of Trajectory, an AI research and product company building a continual-learning platform for enterprise AI systems. Previously an AI researcher at Windsurf, where he worked on SWE-1, and later worked at Google DeepMind.
Long Tool-Use Trajectories Expose Self-Distillation’s Stability Limits
Ronak Malde of Trajectory argues that on-policy self-distillation can make production agent traces usable as dense training signal: a model learns from its own trajectories by matching a version of itself given privileged hints. He says the method avoids the parallel rollouts and trajectory-level rewards of GRPO, but breaks down on long, tool-using tasks when the teacher repeatedly corrects a student that has drifted off course, producing what he calls the “but wait” problem. Malde’s proposed remedies—step-level divergence weighting and residual guidance—are meant to preserve useful correction without teaching the model to hedge or exploit answer-revealing hints.
AI Value Is Shifting From Models to Operating-Layer Control
AI is shifting value toward those who control the layer beneath the interface: iOS permissions and user context, enterprise token flows, compute capacity, data centres and ownership accounts. John Gruber argued that Apple’s AI test is not lateness but whether it will let third-party agents operate deeply inside iOS, while Brad Gerstner argued that enterprise AI spending can keep growing through optimization because tokens and physical infrastructure remain scarce. Kyle Kuzma’s investing comments fit the same ownership frame, treating athlete access as a way to build long-term stakes beyond basketball.