Orply.

Laurie Voss

Head of Developer Relations at Arize AI, where he teaches developers how to evaluate and improve AI applications. He previously led developer relations at LlamaIndex and was a cofounder and founding CTO of npm, Inc.

Longer Skills Files Shift Agent Design From Compression to Verification

Laurie Voss, head of developer relations at Arize AI, argues that frontier models’ capacity to retain simple prompt instructions has risen from roughly 200–300 constraints to thousands in about a year. Re-running and extending the IFScale benchmark, he found that the old compression problem has largely receded, but reliable compliance has not: models now fail through omissions, safety refusals, exhausted reasoning budgets or polished-looking partial answers. His practical conclusion is that teams should verify requirements in outputs rather than assume a fluent response followed the prompt.

AI EngineerSep 9, 20268 min read

Choosing The Right Eval Matters More Than Tuning The Judge

Laurie Voss of Arize argues that agentic applications need the same engineering discipline as other production software: instrumentation, inspectable traces, targeted evals, and controlled experiments, not a handful of prompts that “look right.” In a hands-on workshop using a financial analysis agent, Voss shows how teams should read traces before writing evals, classify failures by root cause, and combine deterministic checks, LLM judges, custom rubrics, and human-labeled meta-evaluation. His central warning is that the choice of eval can dominate the result: the same agent scored 0 out of 13 on a correctness eval and 13 out of 13 on a faithfulness eval because the first judge was asking the wrong question.

AI EngineerMay 14, 202624 min read