AI news · October 2, 2026

Ataraxos beats a top Stratego player with a small training budget

researchreinforcement-learninggamesagents

Made With Models illustration for this story

Nature published a paper on September 30 describing Ataraxos, an AI for games with hidden information. In a 20-game series it beat Pim Niemeijer, a four-time Stratego world champion, 15-1 with four draws. The system used self-play, a belief network to estimate concealed pieces, and search at decision time; the authors report a training cost of a few thousand dollars.

For builders, the result is a pattern for tasks where the system cannot see the whole state. A separate model can represent missing information, while a policy model proposes actions and a search step tests them. It is a pattern, not proof it works in business.

Look for a bounded simulator or replay set in your own problem. Measure hidden-state guesses, action quality, compute cost, and failure cases separately, then test whether a small search step improves accepted outcomes enough to justify the extra latency.

Source: Nature ↗ — Made With Models writes the brief; the reporting is theirs.