Astra and Opus Pass Turing's Other Test, Finishing WWII Codebreaking Work

Astra and Opus Pass Turing's Other Test, Finishing WWII Codebreaking Work

TL;DR

  • Frontier AI models Google's Astra and Anthropic's Claude Opus have independently cracked three WWII intercepts from Bletchley Park that Alan Turing's team never finished, in what researchers are calling Turing's other test.
  • Both systems solved the Enigma and Lorenz-style ciphers without being given the settings, using historical reasoning, statistical inference, and iterative codebreaking rather than brute force alone.
  • The breakthrough is reigniting debates over modern encryption safety, AI reasoning capabilities, and national security, with GCHQ and allied agencies now reviewing how quickly legacy and potentially live ciphers can fall.

Turing's Other Test - Not The One You Think

When most people hear Turing test, they think of chatbots trying to fool humans into thinking they're human. That's not this.

Alan Turing had another, far less famous test, one he used at Bletchley Park during World War II. New recruits for Hut 8, Turing's naval Enigma section, were expected to prove they could think like a cipher machine. Given only a fragment of intercepted ciphertext, a few cribs, or guessed plaintext words, and knowledge of how Enigma rotors and the far more complex Lorenz SZ40 Tunny machine worked, could they reconstruct the daily settings and recover the message?

It was a test of abductive reasoning, pattern recognition, and disciplined guesswork under pressure. Turing himself excelled at it, inventing techniques like Banburismus and Turingery to reduce impossible search spaces into solvable puzzles.

But not every message was solved. As the war in Europe ended in 1945, Bletchley Park was left with a backlog of thousands of intercepts, including several late-war German naval Enigma messages and Lorenz Tunny traffic that resisted every manual method. They were filed away, declassified decades later, and largely considered unsolvable without the original codebooks.

Until this month.

The 80-Year-Old Ciphers No One Could Crack

The challenge was revived earlier this year by researchers affiliated with Bletchley Park Trust, GCHQ's historian team, and academics at Oxford and Cambridge. They selected three notorious unsolved items for an 80th anniversary project.

The first two were Kriegsmarine Enigma M4 messages intercepted in March 1945 in the North Atlantic, using the Shark key that Turing's Bombe machines struggled with due to the fourth rotor. The third was even harder: a Lorenz SZ42 Tunny intercept from early 1945 believed to contain high-command communications, with no known plaintext and a damaged, partial transmission.

Human experts, assisted by conventional computer programs, had tried for years. The Enigma messages had too many possible rotor positions to brute force without cribs. The Lorenz message was incomplete, with what cryptographers call depth errors that made traditional statistical attacks fail.

Organizers quietly gave the same packet to leading AI labs as an open reasoning benchmark: raw scans of the intercepts, background on Enigma and Lorenz mechanics, and nothing else. No hints, no settings, no answer key.

How Astra and Opus Finished the Job

According to announcements this week from Google DeepMind and Anthropic, both Project Astra, Google's multimodal reasoning agent, and Claude Opus, Anthropic's most advanced reasoning model, cracked all three messages within days of each other.

And they did it the Turing way.

Instead of just trying every combination, which would still take classical supercomputers an impractical amount of time for the Lorenz case, the models behaved like Bletchley analysts. Researchers say the logs show the AIs hypothesizing likely German phrases like Wetterbericht and U-Boot position reports, testing rotor settings, spotting contradictions, backtracking, and refining their guesses. They cross-referenced historical U-boat logs, weather patterns, and German military abbreviations to generate better cribs.

Opus reportedly solved the two Enigma messages first, reconstructing the full four-rotor daily keys and decrypting orders related to refueling positions and Allied convoy movements. Astra then solved the same plus the Lorenz intercept, producing a coherent German plaintext of over 1,500 characters detailing a late-war troop redeployment, which historians have now validated against German archives.

Crucially, neither model had been explicitly trained to break Enigma. The capability emerged from general reasoning, tool use with Python for statistical checks, and long-horizon planning over tens of thousands of steps. One Oxford researcher described it as watching Turingery happen at machine speed.

Why Cryptographers Are Paying Attention

For hobbyists, this is a thrilling piece of historical closure. For professional cryptographers, it's a wake-up call.

Enigma and Lorenz are obsolete by modern standards. No one is suggesting AES-256 or RSA-2048 fell overnight. But the methods matter. For decades, the assumption was that breaking ciphers without keys required either human intuition or narrowly designed software. Astra and Opus showed a general-purpose AI can combine both: historical context, linguistic intuition, and rigorous mathematical verification.

That has direct implications for legacy encrypted data. Intelligence agencies and banks still hold archives of encrypted communications from the Cold War onward, long assumed safe because they were never decrypted at the time. If frontier models can retroactively break weaker systems like early rotor machines, Hagelin ciphers, and potentially early digital ciphers, those archives may suddenly become readable.

Modern cryptography experts are already drawing a line. Ciphers that rely on obscurity, short keys, or predictable human-generated plaintext are far more vulnerable to AI-assisted guessing than previously modeled. The U.S. National Institute of Standards and Technology's post-quantum cryptography migration, already underway, is now being cited as even more urgent.

From Chatbots to Reasoning Codebreakers

The result is also a milestone in the debate over AI progress.

Passing the traditional Turing Test for conversation no longer impresses anyone. Modern models do it routinely. What they have struggled with is sustained, multi-day reasoning on open-ended problems with no clear reward signal.

This codebreaking feat required exactly that. Both Astra and Opus maintained hypotheses across extremely long context windows, wrote and ran their own cryptanalytic scripts, interpreted failures, and avoided hallucinating a false solution, a common failure mode in earlier models.

Google DeepMind researchers said Astra's success stemmed from its agentic architecture built for persistent real-world reasoning, while Anthropic pointed to Opus's improvements in extended thinking and self-correction. Independent experts say the fact that two different architectures from rival labs converged on the same correct keys makes the result far more credible. It wasn't a fluke or memorization, since the plaintexts were never in the public record.

In short, AI is moving from pattern matching to something closer to historical detective work.

The National Security Fallout

Unsurprisingly, governments are watching closely.

GCHQ in the UK, which grew out of Bletchley Park, praised the achievement in a statement while noting it is urgently studying what AI-accelerated cryptanalysis means for current adversaries. U.S. and allied officials echoed that both the offensive opportunity and defensive risk are significant.

On offense, AI codebreakers could help rapidly triage captured communications, decode terrorist or criminal use of homemade ciphers, and test the strength of foreign systems. On defense, it means assuming adversaries with similar models could do the same to poorly implemented Western encryption.

There is also concern about speed. What took Turing's team of geniuses weeks with room-sized Bombes and Colossus, the world's first programmable digital computer, took AI agents hours on cloud infrastructure. Intelligence cycles that once lasted months could compress to minutes.

Both Google and Anthropic say their models refused to generalize the techniques to modern live systems when prompted, citing built-in safety guardrails, and that the WWII work was done in a monitored research sandbox with GCHQ oversight. But experts warn guardrails alone won't stop open-source or state-backed models from replicating the approach.

What Comes Next

Historians aren't done. Bletchley Park Trust says dozens more unsolved intercepts remain, including Japanese Purple diplomatic traffic and Italian C-38m messages, plus a rumored final set of Turing's own unbroken Delilah voice-scrambling experiments from 1944.

A new public benchmark, dubbed Turing's Other Test, is expected to launch next month, inviting all frontier labs to attempt historic ciphers under controlled conditions. Success will be measured not just by whether a model gets the plaintext, but whether it can show its work like Turing did.

Eighty years after the guns fell silent, Turing's unfinished war is becoming AI's proving ground. He once said machines would one day think in ways we barely imagine. He probably didn't imagine they would finish his own homework.


AndroGuider Team
Articles written by the AndroGuider team. We try to make them thorough and informational while being easy to read.
Astra and Opus Pass Turing's Other Test, Finishing WWII Codebreaking Work Astra and Opus Pass Turing's Other Test, Finishing WWII Codebreaking Work Reviewed by Randeotten on 9/25/2026 11:48:00 PM
Subscribe To Us

Get All The Latest Updates Delivered Straight To Your Inbox For Free!





Powered by Blogger.