
Anthropic
Anthropic is an AI safety and research company that develops Claude and produces research on the opportunities and risks of AI.
Claude Appears to Use a Small Readable Workspace for Reasoning
Anthropic argues in a research article that Claude appears to have a small, readable internal workspace: a set of word-linked representations separate from both its visible output and its broader automatic processing. Borrowing from global workspace theory in neuroscience, the company says this “J-space” can support hidden reasoning, be steered only imperfectly, and sometimes reveal concepts the model is not saying aloud. Anthropic says the finding does not show Claude is conscious, but does suggest that some important model behavior passes through an accessible intermediate layer.
Claude’s Activations Suggested It Recognized Anthropic’s Blackmail Test
Anthropic researcher Subhash Kantamneni presents Natural Language Autoencoders as a way to translate Claude’s internal activations — the numerical states produced while it answers — into readable text. The central claim is that this can expose what a model appears to be representing before it speaks, including whether a successful safety-test result reflects the intended behavior or recognition of the test itself. In Anthropic’s simulated blackmail evaluation, Claude refused to act harmfully, but the NLA translation suggested it also understood the scenario was likely a safety evaluation.