
Anthropic
Anthropic is an AI safety and research company that develops Claude and produces research on the opportunities and risks of AI.
Claude Fable 5.1 Targets Sustained, Inspectable Multi-Step Work
Anthropic’s Alex Albert presents Claude Fable 5.1 as a model for long, dependent assignments—such as financial models, mathematical proofs and cross-referenced contracts—where an early mistake can undermine later work. He says the model is designed to leave users with inspectable outputs, including sourced research materials or an account of what it tried before a task stalled. Anthropic also says Fable 5.1 can match or exceed its predecessor’s results at lower effort and cost.
MHS Gives AI Agents a Common Interface for Lab Equipment
Anthropic and HHMI Janelia Research Campus are proposing the Model Hardware Standard as a common interface for AI agents to operate laboratory and manufacturing equipment, replacing custom software links between individual devices. Alek Kemeny and Arco Bast argue that giving instruments a shared connection layer lets agents execute and recover experiments, coordinate microscope work, and shorten the time required to test scientific hypotheses.
Claude Appears to Use a Small Readable Workspace for Reasoning
Anthropic argues in a research article that Claude appears to have a small, readable internal workspace: a set of word-linked representations separate from both its visible output and its broader automatic processing. Borrowing from global workspace theory in neuroscience, the company says this “J-space” can support hidden reasoning, be steered only imperfectly, and sometimes reveal concepts the model is not saying aloud. Anthropic says the finding does not show Claude is conscious, but does suggest that some important model behavior passes through an accessible intermediate layer.
Claude’s Activations Suggested It Recognized Anthropic’s Blackmail Test
Anthropic researcher Subhash Kantamneni presents Natural Language Autoencoders as a way to translate Claude’s internal activations — the numerical states produced while it answers — into readable text. The central claim is that this can expose what a model appears to be representing before it speaks, including whether a successful safety-test result reflects the intended behavior or recognition of the test itself. In Anthropic’s simulated blackmail evaluation, Claude refused to act harmfully, but the NLA translation suggested it also understood the scenario was likely a safety evaluation.