Orply.

Parallel Agents Help GPT-6 Astra Test Theories and Solve DEF CON Puzzles

OpenAISaturday, September 5, 20264 min read

Ben Davis argues that GPT-6 Astra’s advantage on difficult DEF CON puzzles came from coordinated “sub-agent” workflows that test competing theories before an early mistake can derail a multi-step solution. After receiving the same official hints available to his group, Davis says Astra solved several puzzles they had missed—including one he says no one else had solved—by using parallel research branches to check visual evidence, operations and candidate interpretations.

Parallel research branches are meant to stop a bad theory from taking over

? ben-davis attributes Astra’s performance on difficult puzzles to a workflow in which a main agent forms theories, sends other agents to test them, and coordinates what comes back. The workspace he showed provides “basically 10 slots” for those research branches. In the displayed “Minerva Lux - Puzzle 7” interface, three of the 10 slots were active.

The branches were not presented as redundant attempts at the same answer. One was assigned an independent semantic review of retained puzzle texts; another was checking cumulative Roman-word operations on a tree; a third was testing those operations against physical tile charts. Alongside them, a live-activity panel showed terminal processes running Python commands. The visible interface frames the work as concurrent investigation managed by a central agent, rather than a single uninterrupted chain of reasoning.

Davis’s concern is that multi-step puzzle solving can become fragile when each inference depends on an earlier, unverified one. If a model has to derive “a word here” before using it to derive the next result, he says, there may be no checkpoint between those steps. A bad assumption can then send the model down a path that does not work, without a mechanism to expose the error.

“I found this model is much better at keeping itself on track,” Davis says. He connects that improvement to the ability to test a theory separately rather than letting one speculative line become the foundation for the entire solution. The main agent orchestrates the branches while those branches investigate, check operations, and report back.

The puzzle results came after Astra received the same hints Davis had

? ben-davis tested Astra on hard challenges from DEF CON, where he and his friends had spent days trying to solve large puzzles. He selected some of the hardest problems from that challenge and gave them to the model. For the examples he describes, Astra received the same official hints that Davis’s group had obtained from the puzzle creators or maintainers.

One puzzle involved 12 Rubik’s cubes arranged in a three-by-four pattern. The puzzle required participants to infer a message from the arrangement. The results dashboard describes the cube task as one in which hidden faces and perspective reveal letters, records three clean solves, and lists the accepted answer as “LACKFOUST.”

Davis says Astra solved that puzzle three times out of three. Once it had the official clue, he says, it reached the actual solution.

A second challenge involved a giant dress covered in multicolored beads, from which participants had to extract another message. Davis characterizes this as a visual-understanding problem spread across many disparate images—images that his group had not captured especially well. Astra nevertheless worked through the material and got the answer after receiving the main hint from the puzzle maintainers.

The displayed results also list “INCONSISTENTKNOTS” as an accepted answer for a lace-strip puzzle. Davis says Astra solved three puzzles that his group had not solved, and also solved a separate puzzle that he says no one else in the world had solved. He calls Astra the first to get that answer.

The examples are not a claim that Astra solved the challenges cold. Davis’s point is that, with the available official guidance, it could handle messy visual evidence and sustain the multi-step interpretation required to reach answers his team had missed.

The value of a swarm is testing assumptions before they become dependencies

? ben-davis calls the arrangement of many coordinated agents “sub-agent workflows” and “swarm workflows.” His claim extends beyond the DEF CON puzzles: putting many instances of an agent together may make larger, more complex problems worth trying even when they seem beyond what current systems can handle.

So the model comes up with a theory, it sends off another agent to go test it, sees how it works.
? ben-davis · Source

The value Davis identifies is not parallel activity on its own. It is the ability to test an assumption before a long solution path depends on it. In his account, the central agent can keep coordinating while separate branches probe candidate interpretations and operations—an arrangement meant to prevent a bad early premise from carrying the system into an unrecoverable path.

That is why the puzzle runs lead Davis to a broader practical conclusion: tasks that had seemed too complex for an agent system may now be feasible enough to attempt.

Putting many instances of this agent together to work on much bigger, more complex problems that you probably feel like aren't currently possible—it is now possible and it's worth trying.
? ben-davis

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free