An AI system named Ataraxos has defeated Jeroen Niemeijer, widely regarded as the greatest Stratego player of all time, marking a significant milestone in the development of artificial intelligence for games with imperfect information. The victory suggests that strategic decision-making in complex, hidden-information environments is no longer exclusively the domain of human intuition.
What Happened
Niemeijer, who has won four world championships and spent over 600 weeks at the top of the rankings, faced Ataraxos in a series designed to test the limits of current AI capabilities. According to the source, Stratego is exceptionally difficult for AI because players set up 40 pieces face down, creating more than 10^33 possible setups. In such games, the value of a move depends on prior hidden information, a challenge that has historically kept humans superior.
The series against Niemeijer was spread over three weeks to allow for preparation. Niemeijer was incentivized with a $1,000 participation fee plus $100 per win and $50 per draw. The authors report that Ataraxos achieved an effective win rate of 85 percent, counting draws as half a win. This margin is described as unprecedented at the highest level, particularly because the AI operated at a structural disadvantage: Niemeijer could adapt his style over the course of the series, while Ataraxos could not adapt to him.
Why It Matters
The victory is notable not just for the outcome, but for the efficiency of the training process. DeepMind’s previous attempt, DeepNash, failed to surpass top human players and reportedly required 1,024 TPU nodes for two to three months, with estimated costs between $3 million and $4.5 million. In contrast, the Ataraxos team utilized a university budget, training on 16 Nvidia H100 GPUs for one week, plus four additional days on four GPUs for its belief network. The researchers estimate the total cost at less than $8,000, roughly 1/500th of the compute cost of DeepNash.
The key to this efficiency, according to first author Samuel Sokota, was a combination of higher sample efficiency and a custom GPU simulator. The system learns purely through self-play without human data. It employs a regularization technique that forces the AI to vary its play early in training to avoid getting stuck in fixed strategies, then gradually reduces this pressure to allow for more deliberate, small corrections. This approach helps the AI handle the randomness required in Stratego without making the learning process chaotic.
The implications extend beyond board games. The researchers argue that large amounts of hidden information are no longer an obstacle for reinforcement learning. They point to financial markets, military conflicts, and negotiations as practical domains where this technology could be applied, provided fast and accurate simulators can be built. The system also demonstrated versatility by setting new records in the cooperative card game Hanabi and beating previous benchmarks in the Chinese card game Dou dizhu.
The Bottom Line
Ataraxos has successfully ended the era of human superiority in high-level Stratego, achieving this feat with a fraction of the resources required by previous industry leaders. While the researchers acknowledge that the current search method has a performance ceiling that cannot be lifted simply by adding more compute time, the public availability of the code and the demonstrated efficiency in imperfect information games signal a shift in how AI approaches complex strategic problems.