Hackathons

One so far, and it taught more about measurement than about chess.

AI Chessathon, online qualification, 4 to 11 September 2026

A chess engine in pure Python, measured before it shipped

Every agent runs as plain Python on one CPU core with 2 GB, 120 seconds plus half a second a move, no native code and a 90-second import budget. Twenty-four builds over eight days, twenty-one of them uploaded and validated. The ladder closed after 109 rated rounds at 47th of 465, top 10 percent, rating 2319 and a best rank of 7th. Builds were then locked for a 13-round Swiss, which finished 28th of 334 on 9.0 out of 13.

The transferable part was not the engine but the harness around it: an A/B rig with Elo confidence intervals, and a rule that no change ships without a number. Around fifty ideas were rejected with the figure that killed each one, including several that shipped early, lost games on the ladder, and were reverted.

Magic bitboards, principal-variation search, residual NNUE, own-engine opening book, numba

Builds
24
Best ladder rank
7th
Ladder close, of 465
47th
Final Swiss, of 334
28th

The final qualification Swiss

On the afternoon of 11 September the ladder closed, every team's build was locked, and the 334 teams that entered played a 13-round Swiss.

Thirteen rounds, 9.0 points 7 wins 4 draws 2 losses Where that finished, in each field Qualifier ladder 47th of 465 465th Final Swiss 28th of 334 334th
Nine points from thirteen rounds, and where that sat in the field. The Swiss was the engine's strongest stage: a narrower field than the ladder and a better placing in it.

I finished 28th of 334 on 9.0 out of 13, at a stage rating of 2516, which was the engine's best rating of the event. It played the whole Swiss on v24, the build with the deep opening book.

That was in range of a seat at the London final on 12 September, where the room holds 50 and seats go one per UK university student down the Swiss standings. I could not attend on the day, so I withdrew, and round 122 was the engine's last competitive game. Across both stages it finished 42 wins, 41 draws and 33 losses in 116 decided games.

Swiss score
9.0/13
Record
7W 4D 2L
Stage rating
2516
Both stages
116 games

What the engine is

Board. Magic bitboards, copy-make, Zobrist hashing, incremental material and piece-square sums.

Search. Principal-variation search with a packed transposition table, aspiration windows, null-move pruning, late-move reductions, reverse and forward futility, static-exchange pruning, singular extensions, killer, counter and history ordering, check extensions, and a quiescence search with check evasions.

Evaluation. A residual NNUE with 768 piece-square inputs per perspective, 256 hidden units and three output buckets by piece count, trained on about 7 million Stockfish-labelled positions on top of a hand-tuned tapered evaluation. Accumulators update incrementally in one fused pass per move.

Clock. Remaining time spread over max(30, 60 minus half the ply) moves plus 60 percent of the increment, with a hard ceiling of half the clock. An opening book of 858 positions, every start and opening line the engine actually met in its own rated games, each answered by the engine itself after 30 to 40 seconds of thought. A draw is scored at minus 60 centipawns so it plays on in level positions.

Roughly 600 to 700 thousand nodes per second on one core, reaching depth 10 to 12 in the two to three seconds the clock allows.

Build log highlights

Nothing shipped without beating the previous build in a measured match.

0+50 Elo+100 Elo v2+75Static exchange evaluationv3+109Pondering on the opponent's clockv10 to v12+37First NNUEv14+243.4M-position net, mirror augmentationv17+8135 percent more node speedv19+54Clock scheme restoredv21+4525 percent more node speedv22+206.95M-position netv23+10Contempt 60, relabelled rows
Elo gain per shipped build, measured against the previous build in 300 to 1,800 games. Whiskers are 95 percent intervals where they were recorded; v22 is a range, v23 a net figure. Pondering was banned by a rules change after v3, so that +109 did not last. v24, the deep book, is safe by construction and has no Elo number.

What did not help, each measured

Recorded so the same idea is never tried twice on a hunch.

Where the errors actually were

Every game on the ladder was reviewed with Stockfish after the round. By day three the endgame was already level with the top three, and everything left was judgement in the opening and middlegame.

Opening, first 8 1.4 4.9 Middlegame 2.3 4.1 Endgame 1.7 1.5 02.55
Top three teamsThis engine
Inaccuracies, mistakes and blunders per 100 of the engine's own moves, from the site's own reviews of 833 games across the top 30 teams after day three. Lower is better.

The honest read. The field improved faster than the engine did. The rating held between 2200 and 2372 through the last three days, so the drift down the table was the field growing from 400 teams to 465, not the engine getting worse. Losses were rarely tactical. They were quiet positional slides of 40 to 50 centipawns a move from around move 10 in closed structures, which is an evaluation problem, not a search one, and eight days was not long enough to fix it.

The one time the rule lapsed, three clock changes shipped on game evidence rather than a measured match, the rank fell from 7th to 66th in a day. Measuring them properly showed they were harmful, v19 reverted them, and the rank was back to 17th by the end of the next day. The record of what did not work is the useful output, and the same discipline now runs the ball-on-plate build: comparison apparatus first, then the thing being compared.