# GPT-6 Astra Cracked an 82-Letter Enigma Message. The Case Study Says a Researcher Led.

**Summary:** The 82-letter MVUEH plaintext and key survive a separate re-decryption. I accept the break, but the split between 'entirely on its own' and 'researcher-led' is why agent claims need a second audit trail.

- Canonical: https://markhuang.ai/news/gpt-6-astra-enigma-researcher-led
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-22
- Section: News
- Tags: GPT-6 Astra, Enigma, AI Agents, Cryptanalysis, Research Verification
- Source: [Crypto Cellar Research](https://www.cryptocellar.org/bgac/the-mvueh-break.html)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![An archival Enigma machine sits between a researcher's hands and a branching computational signal](https://cdn.markhuang.ai/news/gpt-6-astra-enigma-researcher-led/hero.webp)

*The plaintext is checkable. Reconstructing who steered the search takes a different kind of evidence.*

Frode Weierud's [Crypto Cellar account](https://www.cryptocellar.org/bgac/the-mvueh-break.html) says GPT-6 Astra helped recover MVUEH, an 82-letter German Army Enigma message sent on July 10, 1941 and left unsolved in the collection since 2005. The result is unusually concrete for an AI capability story: a published key turns the archived ciphertext into connected German, and the corpus curator accepts the break.

I accept the break too. I do not accept the cleanest version of the autonomy story yet. Crypto Cellar says Astra did the work "entirely on its own," while the solver's detailed exhibit calls it a "researcher-led investigation" in which the researcher set the goal and pushed the work forward. That disagreement does not weaken the recovered plaintext. It changes what this episode proves about agents.

## The plaintext is easier to verify than the workflow

The winning search borrowed a 14-letter crib, `ROSENOWROSENOW`, from SIPVX, a related message that Alex Shovkoplyas broke in 2017. A crib is a suspected piece of plaintext. It narrows the search rather than supplying the whole answer. In this case, the other 68 letters were free to become noise. Instead, they formed a request for a route of march and an immediate radio reply.

The solver's [published case study](https://mvueh-enigma-solved.carterl.chatgpt.site/) provides the machine settings, source-letter alternatives, search code, and reproduction packages. It says the recorded message header also selects the recovered physical settings from 26 body-equivalent possibilities. Those are much better receipts than a screenshot of a fluent answer.

A [separate verification repository](https://github.com/swarm-ai-research/cipher-break-verification) starts with Crypto Cellar's published ciphertext rather than the solver's code. Its small Enigma implementation produces connected German, with garbling in 8 of the 82 positions. Moving one ring setting makes the recognizable phrases disappear. The verifier did not repeat the entire discovery search or prove that the key is unique, and it says so plainly.

> **Info:**
>
> The key and plaintext have a reproducible check outside the model's own narration. The full path that found them, including every human intervention, is a separate claim.

## The two accounts assign the steering differently

Weierud writes that Carter Leffer asked Astra to try the unbroken messages on Crypto Cellar. In that account, Astra selected MVUEH, connected it to SIPVX, chose the Rosenow crib, and built Python and C++ tools for an Enigma simulator and Bombe search. Weierud also says the team was still examining the logs when he published the page.

Leffer's case study draws the boundary differently. It says a researcher led the investigation with GPT-6 Astra and parallel specialist agents, set the goal, and pushed the investigation forward. Agents examined sources, wrote search programs, ran experiments, and reviewed the result. The work spanned September 14 and 15, 2026, but the case study says its logs do not provide a complete total for human or model reasoning hours.

I find the detailed account more useful because it does not turn a collaborative investigation into a magic trick. A human can stay quiet for long stretches while still defining the problem, approving a new direction, or deciding when the evidence is good enough. "Entirely on its own" erases those distinctions. "Researcher-led" leaves room to inspect them.

This is not a semantic complaint. If I am evaluating an agent for research, incident response, or software work, I need to know whether the model found the decisive connection unprompted, chose among human suggestions, or continued after a researcher rejected failed approaches. Each version implies a different level of supervision and a different staffing plan.

## The published workflow still earns attention

The case is still strong. The search dealt with uncertain handwriting, competing transcriptions, failed hill-climbing controls, rotor behavior, plugboard settings, and an independent header. The case study reports 43,016 completed search batches inside its declared assumptions and about 4.29 billion combinations of rotor behavior and crib placement. It also publishes unsuccessful experiments instead of presenting the winning path as inevitable.

That is close to the kind of multistep work OpenAI advertises for [GPT-6 Astra](https://openai.com/index/gpt-6-astra/): browsing, coding, scientific investigation, and sustained tool use. It is not an OpenAI evaluation, though. It is one researcher-run case with an unusually checkable output. I would treat it as evidence that Astra can participate in a demanding technical search, not as a general benchmark for autonomous research.

Public discussion shows how quickly that distinction disappears. A [Reddit thread asked whether Astra could now solve the remaining Zodiac ciphers](https://www.reddit.com/r/ZodiacKiller/comments/1wjbbm5/if_gpt6_astra_can_crack_unsolved_german_enigma/). One reply pointed out the missing ingredient: some short ciphers may not contain enough information to identify a unique answer. More compute cannot manufacture evidence the ciphertext never contained. MVUEH worked because the related message, alternate source copy, recorded header, and plausible crib constrained the problem.

## Agent claims need two audit trails

For this kind of announcement, I want one audit trail for the result and another for agency. The first should include inputs, code, candidate outputs, and a way to reproduce the answer. The second should record the initial instruction, human steering, agent delegation, model and tool versions, checkpoints, rejected paths, and the conditions that ended the search.

That second trail matters for the same reason the harness mattered in my earlier look at [Astra's 45.1-point ARC score gap](https://markhuang.ai/news/astra-arc-score-45-1-point-harness-gap). A model name is not a complete description of a working system. Here, the parallel agents, researcher decisions, source archive, simulators, and search budget all belong to the result.

My label would be precise: this is a verified Enigma break produced in a researcher-led investigation with GPT-6 Astra and parallel agents. The claim that Astra independently chose and completed the whole project needs a trajectory that reconciles the two published accounts. For now, I trust the decryption. I would not use it as proof that Astra ran an independent research project.
