For immediate releaseHutter Prize Committee

Energy-Limited, Leak-Proof AI Benchmark Posts Its Biggest Gain in 20 Years as Trillions Ride on AI

The Hutter Prize requires no validation data, so its scores cannot be inflated by data leaks. Three entries, each running on a single CPU core, just improved its record by nearly 10%.

— With trillions of dollars in data-center and power investment riding on claims of AI progress, the Hutter Prize for Lossless Compression of Human Knowledge today awarded the largest advance in its 20-year history on a benchmark that data leaks cannot inflate. Three entries together improved the record by nearly 10%, compressing a 1GB snapshot of Wikipedia by more than a factor of 10, each on one general-purpose CPU core.

Line chart of Hutter Prize compression factor from 2006 to 2026. The enwik8 track, normalized to 1GB, rises from about 6.6 in 2006 to about 8.6 in late 2017. The enwik9 track continues from about 8.6 in 2019 to about 9.2 in mid-2026, then jumps to about 10.3 with the 2026 awards.
Hutter Prize record compression factor, 2006–2026.1 Source: Hutter Prize committee.

Most AI benchmarks score models on validation data: material held back from contestants so models cannot simply memorize it. If that validation data leaks into a model’s training data, by accident or otherwise, its score rises with no real gain in capability, and outsiders have no easy way to tell. The Hutter Prize requires no validation data. The full 1GB file is public, and entries are scored on how small they can make it, counting the program that rebuilds it. Anything a model memorizes has to be stored in that program and counts against its score.

The prize measures intelligence the way science tests a theory: by how well it predicts. Every bit a model fails to predict is a bit that must be stored, so the compressed file is a theory of human knowledge that pays for each of its mistakes.

“This year has seen the largest progress in its 20-year history,” said Marcus Hutter, the prize’s sponsor, author of the 2026 book Job-Less Utopia and advisor to Google DeepMind co-founder Shane Legg’s 2008 PhD thesis, “Machine Super Intelligence.”

Fixed compute, rising capability

Entries must run on a single general-purpose CPU core, within roughly 50 hours and under 10GB of RAM. Because the hardware budget is held constant, every gain on the leaderboard comes from better algorithms rather than more chips or more electricity. The result is a counterpoint to the assumption that progress in machine intelligence must track growth in data-center power demand. The committee sets these limits for three reasons:

  1. Specialized hardware biases research toward the industry’s current assumptions.
  2. Talented individuals, not only well-funded labs, can afford to compete.
  3. The public increasingly opposes the resource consumption of superintelligence efforts.

The committee likens resource-limited compression to decision accuracy under a reaction-time limit. Unlike bell-curve measures of intelligence, it sits on an absolute scale, and rote memorization is penalized by its size.

Next-token prediction, a decade early

The prize launched in 2006 on the premise that predicting text is the core of intelligence, anticipating today’s next-token-prediction language models by more than a decade. The committee draws a sharp line between this benchmark and the Turing Test: the Turing Test checks only whether a machine can fool a human, while the Hutter Prize tests for superintelligence.

The winners

Each improvement is measured against the record set by the entry before it. The three awards total nearly €50,000.

Judging took longer than usual because the rate of new entries has exploded. Contestants are now using superintelligence to improve their machine learning algorithms, which caught the judging committee by surprise; it is now racing to use the same tools to automate judging. This is feasible because the award criteria were designed to leave little discretion to the judges.

The next target

Awards are 1% of €500,000 for every 1% improvement, with a minimum improvement of 1% over the current record. Rules, records and past winners are at prize.hutter1.net.

About the Hutter Prize

The Hutter Prize rewards advances in lossless compression of human knowledge as a rigorous measure of superintelligence. It launched in 2006 at €50,000 scale on enwik8, a 100MB Wikipedia snapshot, and scored entries on compression ratio alone. In 2020 it expanded to €500,000 scale, moved to the 1GB enwik9 snapshot, and began counting the size of the compressor along with the compressed data. The committee is Marcus Hutter (sponsor and arbiter), Matt Mahoney (competition operations) and Jim Bowery (claim verification and public relations).

Media contact: Jim Bowery, Hutter Prize committee
jabowery@gmail.com
prize.hutter1.net

###