All articles
·50 min read

I trained an AI model from scratch in a weekend, and it learned chess on its own

LLMPyTorchTransformerGPUChess
Source code

That weekend, I wanted to dig into AI in some way other than consuming existing models. I use them every day, I do fine-tuning, I run on GPUs. But between "knowing how to use a model" and "knowing how a model learns" lies a gap that only training one from scratch will close.

Not fine-tuning. Not a well-crafted prompt on an existing model. From scratch: empty architecture, random weights, one file of raw data. And a single gaming graphics card.

What follows is the full account of this project: the decisions, the reasons behind them, and the mistakes. Every command is here, in the order I ran it.

The technical is explained as the text goes along, and a full glossary at the end of the article covers every term (, , , , …) without the jargon. No prior knowledge of machine learning is needed to read what follows.

Finding the idea: why chess

I wanted an experiment that answered a sharp question, not a project that produced a vaguely impressive demo. The question I set myself was this: how far can a system that only learns statistical regularities reconstruct a mechanism that was never described to it?

Chess is the ideal ground, and not for the prestige of the game. What matters is that everything in it is mechanically verifiable. A move is legal or it is not. No human judgment, no ambiguity, no fuzzy metric to argue over. I can measure exactly what the model has understood, across tens of thousands of positions, in a few seconds.

Compare that with text: if I train a model to write, how do I objectively measure that it has "understood" grammar? In chess, the question does not arise.

The second reason is just as decisive: the data source is virtually inexhaustible, and it is free. Every month Lichess publishes a dump of every game played on the platform, openly available. A single month, July 2026, contains 89 million games, or 29 GB compressed. And there has been one dump per month since 2013. In total, several billion games available for download, with no account, no quota, no licence negotiation.

That is a luxury you get almost nowhere else. With text, clean corpora are either paid, or legally murky, or already digested by everyone. Here I could afford to be extremely selective: throw away 87% of what I read and keep only the good-quality games, without ever fearing I would run out of material. A single month was enough to reach the 800 million I was aiming for. If I wanted more, I only had to download the previous months: that is exactly what I did to train a second, larger model, which I will come to later in this article.

This abundance changes the nature of the problem: the question is no longer "where do I find enough data?" but "which data is worth keeping?". That is a far more comfortable position, and a far rarer one than people think.

So I imposed on myself the constraint that makes the experiment clean: the model would have access to no rule of the game. It never sees a board. It does not know that pieces, squares or a king exist. It receives sequences of symbols in notation, e2e4 e7e5 g1f3, and its only task is to predict the next one. Exactly the mechanism of a ChatGPT on text, but in a domain where I can verify everything.

First contact with the machine

Before writing a single line of the model, you need to know what you are playing on. I started with the inventory: GPU, disk, CPU, RAM, versions.

On the compute side, it is comfortable: a 24 GB RTX 3090 and a 12 GB RTX 3060, driver 570, CUDA 12.8. Planning room for the data, on the other hand, is not optional: a single monthly Lichess dump weighs 29 GB compressed, and you need three times that once it is decompressed and encoded.

mkdir -p ~/chess-model && cd ~/chess-model
mkdir -p data checkpoints logs

python3 -m venv venv
PY=venv/bin/python
$PY -m pip install --upgrade pip
$PY -m pip install torch --index-url https://download.pytorch.org/whl/cu126
$PY -m pip install python-chess zstandard numpy tqdm matplotlib

First precaution, automatic for me because I run all day long: force the GPU ordering. nvidia-smi sorts by PCI bus and shows the 3060 at index 0; PyTorch sorts by "fastest card first", so its cuda:0 is the 3090. The two tools give contradictory numbers for the same card. All my scripts therefore force CUDA_DEVICE_ORDER=PCI_BUS_ID, which aligns PyTorch with nvidia-smi. Without it, you think you are training on the big card while you saturate the small one, and nothing warns you.

The graphics card that lies

With both cards freed up, I ran a benchmark. The principle is simple: multiply two large matrices in a loop and time it. An n × n product costs 2n³ operations, so dividing by the time gives a throughput in . That is the reference measurement because 90% of a 's compute is exactly this.

$PY 
[bench] cuda:0 NVIDIA GeForce RTX 3060 (11745 MiB libres) ...
        -> 27.37 TFLOPS bf16 (matrices 8192^2)
[bench] cuda:1 NVIDIA GeForce RTX 3090 (23873 MiB libres) ...
        -> 25.7 TFLOPS bf16 (matrices 8192^2)

The result made me read my code twice.

The big card was slower than the small one. That is absurd: the 3090 carries 82 multiprocessors against 28, that is 2.9 times more compute units.

First reflex, suspect the measurement. I rewrote the benchmark to sustain the load for 30 seconds instead of one, sampling power, clock speed and in parallel every half-second. If it were an artefact, a clock ramping up too slowly for instance, it would show over that duration.

It did not show: under sustained load the gap even widens slightly, 24.98 against 27.53 TFLOPS. The measurement was sound. The problem was elsewhere, and a single command revealed it:

nvidia-smi --query-gpu=name,power.limit,power.default_limit --format=csv
name, power.limit [W], power.default_limit [W]
NVIDIA GeForce RTX 3090, 170.00 W, 420.00 W

It is all there. Limit configured at 170 W, factory limit at 420 W. The card was capped at 40% of its thermal envelope, and the readings under load confirmed it:

                 3090         3060
horloge SM       596 MHz      1966 MHz
puissance        170 W        144 W  (plafond 175 W, non atteint)
température      48 °C        60 °C
TFLOPS bf16      24,98        27,53

The detail that gives a power cap away: the 3090 is pinned at exactly 170 W, minimum and maximum identical across 47 samples. Natural consumption fluctuates; an imposed ceiling does not. And its clock collapses to 596 MHz when its boost is 2100 MHz: it was running at less than a third of its rated speed. This is not thermal, 48 °C under continuous load is cool, and the driver reported no throttling.

The diagnosis checks out arithmetically, which makes it solid: the 3090 has 2.93 times more compute units but runs at 0.303 times the clock of the other card. Product: 0.89. The throughput ratio actually measured is 24.98 / 27.53 = 0.91. The two match within two points. The power cap explains the whole anomaly; there is nothing else to look for.

sudo nvidia-smi -i 1 -pl 280      # ne survit pas à un redémarrage

Why 280 W and not 420. Because the performance gap between the two was marginal. The power/performance curves of Ampere cards are very flat at the top end: the last 140 W buy only a few percent, at the cost of a lot of heat and noise. And this card was going to run under continuous load for hours, with the second training run lasting twenty hours straight. It was summer, in the middle of a heatwave, in a room without air conditioning. The real constraint was no longer throughput, it was temperature: a card that gets too hot ends up lowering its own clock, and you then lose more than you thought you had gained by raising the ceiling. 280 W is the compromise I kept, and it is at that setting that every figure in this article was measured.

The gain, on the other hand, is anything but marginal. Same matrix product, same duration, same method:

                    170 W        280 W       facteur
TFLOPS bf16         24,98        62,86       x2,52
horloge SM          596 MHz      1506 MHz    x2,53
puissance           170 W        279 W       x1,64
température         48 °C        54 °C

The two factors, throughput and clock, match to within 0.01. That is not a coincidence, it is a confirmation: on this kind of compute, throughput is directly proportional to clock speed as long as nothing else is the limit. Power really was the sole bottleneck. And only six degrees more, for two and a half times the throughput.

That is the first lesson of the project: you never plan against a manufacturer's figure. My entire duration planning would have been wrong by a factor of 2.5. And that measured figure of 62.86 TFLOPS is the one I will compare the against later. Comparing against the spec sheet would make no sense.

The vocabulary: one token per move

This is the most structuring design decision of the project, so it deserves a pause. A language model splits text into ; here, I had to decide what a token is.

I chose one move = one token, indivisible. The has 1971 entries: 1968 moves in notation, plus , and .

Why not a BPE, the standard LLM tokenizer

It is the default choice of every language model, and it would have been a bad one here. With a , e2e4 could be split into e2 + e4, or worse into e + 2e + 4. The model would then spend part of its capacity relearning that these fragments form a unit, when I know that unit in advance. Better to hand it over directly.

The benefit is threefold, and I had only anticipated the first one.

A game of 80 is exactly 80 tokens. The context window is reasoned about in moves rather than characters, which makes sizing trivial: 256 tokens of context is 256 moves, far more than a complete game.

Masking illegal moves becomes a simple boolean mask. Since the vocabulary is finite and known in advance, it is enough to set the of impossible moves to minus infinity. That is what makes it possible to evaluate the model without a net (to measure what it has understood) and to make it play with a mask (so that it never produces an illegal move). Two uses, two different questions, the same model.

The vocabulary depends on no corpus. It follows only from the rules of the game, so it is identical for every dataset. It is in fact a source of confusion when you juggle several corpora: I keep it at the root, never in a data subdirectory.

The 1968 is not a chosen number. Rather than trying to guess which moves are legal, which depends entirely on the position, I enumerate every displacement that is geometrically possible on an empty board.

The trick is that every piece moves either like a queen (ranks, files, diagonals) or like a knight. The king, the rook, the bishop and the pawn only make subsets of the queen's moves. So from each of the 64 squares, it is enough to trace the 8 queen rays and the 8 knight jumps, then add the promotions.

The result is a strict superset of the legal moves, and that is exactly what you want: it will never be missing a move, and the model will learn on its own that some never appear. The opposite would have been dangerous. A vocabulary that is too tight makes encoding crash the day a game contains an unforeseen move, and you find out in the middle of an hour-long job.

The list is sorted, therefore deterministic: the same vocabulary rebuilds identically on any machine. That avoids the silent disaster of a trained with a different token ordering, where the model plays nonsense without a single error being raised.

The vocabulary file

The file has 1971 entries and looks like this, with the itos list (index to move), its inverse stoi, and the ids of the three :

{
  "itos": ["<pad>", "<bos>", "<eos>", "a1a2", "a1a3", ..., "h8h7"],
  "stoi": {"<pad>": 0, "<bos>": 1, "<eos>": 2, "a1a2": 3, ...},
  "vocab_size": 1971,
  "n_special": 3,
  "n_moves": 1968,
  "pad_id": 0,
  "bos_id": 1,
  "eos_id": 2
}

The complete file is here: . Click it to open and copy it, then save it as . It is the only data file to fetch in this article, and it is enough on its own: everything else is downloaded or computed.

The data: where I nearly optimised at random

I took the Lichess dump for July 2026.

cd data
wget -c https://database.lichess.org/standard/lichess_db_standard_rated_2026-07.pgn.zst
cd ..

# Vérifier l'intégrité : la taille doit correspondre exactement
stat -c %s data/lichess_db_standard_rated_2026-07.pgn.zst
curl -sI https://database.lichess.org/standard/lichess_db_standard_rated_2026-07.pgn.zst \
  | grep -i content-length

29,051,046,504 bytes, downloaded in 8 min 28 s. Size verified byte for byte.

Why the conversion is expensive

The dump holds the moves in notation, the one used in chess books: Nf3, exd5, O-O. I want : g1f3, e4d5, e1g1.

The difference between the two is the heart of the problem. SAN is a contextual notation: "Nf3" means "the knight goes to f3", without saying which one. If two knights can go there, the notation adds a disambiguator, otherwise it relies on the reader to work it out. UCI is explicit: origin square, destination square, full stop.

Translating from one to the other therefore requires knowing the position, and so replaying the game move after move on a board. SAN is "take the second turning on the right"; UCI is a pair of GPS coordinates. To convert the first into the second, you have to have followed the whole route from the start.

Measure before parallelising

Before writing the final parser, I measured the three cost centres separately. That step takes two minutes and I nearly skipped it.

$PY  benchmark --games 50000
décompression zstd + découpage en parties   130 089 parties/s
filtrage + rejeu SAN -> UCI, 1 cœur            5 163 parties/s
filtrage + rejeu SAN -> UCI, 10 ouvriers      30 619 parties/s
rétention                                        12.15 %

A factor of 25 between decompression and replay. My intuition said exactly the opposite: I thought pulling 200 GB of text out of a compressed file would be the bottleneck, and that was the step I was going to parallelise. I would have optimised the thing that was already twenty-five times too fast.

The architecture follows on its own. A single process decompresses and splits the stream, since it is fast enough by a wide margin; ten workers in parallel replay the games, the only genuinely expensive step. Result: one hour of processing instead of five.

This is the most reusable lesson of the whole project. On a data pipeline, you measure stage by stage before parallelising anything, otherwise you speed up what is already fast.

The lazy generator that nearly swallowed 100 GB of RAM

That parallelisation laid a trap for me, and I am telling it because it is a nasty one: everything looked perfectly reasonable.

Picture a kitchen. I have ten cooks preparing the dishes, that is my parallelisation. And I have a supplier bringing in the ingredients as they go, one crate at a time, precisely so the kitchen is never cluttered. In , that supplier is called a generator.

A word about it, because it is one of the most elegant mechanisms in the language. It is said to be : it computes nothing until something is asked of it. The difference is obvious on two lines that appear to do the same thing:

parties = [traiter(p) for p in dump]      # une liste : tout est calculé, tout tient en mémoire
parties = (traiter(p) for p in dump)      # un générateur : rien n'est calculé avant qu'on le demande

The first line walks through the 92 million games, processes them all, and stores the result in memory before the next line even runs. The second does nothing at all. It merely prepares a recipe, which it will execute one game at a time, as it is iterated over.

Why this is excellent: the memory consumed no longer depends on the size of the data. Processing a thousand games or a hundred million costs the same, since you only ever hold one at a time. So you can work on a file larger than the machine's memory, and processing starts immediately instead of waiting for everything to load. It is the mechanism that makes Python comfortable on large volumes, despite its reputation for slowness.

In short, you trust it blindly. And that is where it goes wrong.

Except I had forgotten about the head of service who hands the work out to the cooks. He waits for nobody. The moment he sees a supplier, he takes everything they have, as fast as possible, and stacks the crates up in the kitchen waiting for someone to deal with them.

That is exactly what Pool.imap_unordered does, the function that spreads tasks across processes. It drains its source at full speed, with no regard for the workers' actual pace. The supplier's caution is therefore worth nothing any more.

The arithmetic is frightening: 92 million games stacked up at once amounts to a hundred or so gigabytes of text, in a machine with 32 GB of memory. The program would have been killed by the system after a few minutes. Probably after I had left, with a log file that stops dead and no explanation.

The fix is simple once you have seen the problem: instead of presenting the whole stock at once, I hand the work out only in batches of fixed size, one after another. The kitchen can no longer be swamped. Measured during the real run: 2.8 GB of memory, stable from beginning to end.

The moral is broader than Python. A cautious component loses all its caution the moment a component in a hurry plugs into it. The part is not at fault, the assembly is, and that is precisely what you do not see when re-reading your own code.

What the filtering throws away, and why

nohup $PY  parse --workers 10 --target-tokens 800e6 \
  --out  --stats logs/phase1_parse_stats.json \
  > logs/parse_full.log 2>&1 &

tail -f logs/parse_full.log

54 minutes later, at 27,307 games/s: 89,288,421 games read, 11,035,777 kept (12.36%), that is 789,834,938 .

So I throw away 87% of the dump, deliberately:

Elo hors de [1800, 2600]   47.0 %      42 318 347 parties
bullet ou ultrabullet      36.8 %      32 376 011 parties
terminaison anormale        3.9 %       3 355 996 parties
moins de 20 demi-coups      0.2 %         201 270 parties
plus de 300 demi-coups                       1 020 parties

Two takeaways. Almost half the games go because at least one player falls outside the 1800-2600 band: the mass of players sits below 1800. And above all, more than a third of the games played on Lichess are bullet, one or two minutes for the entire game. I exclude them because at that time control the moves are dictated by urgency rather than thought: that is noisy data, which would teach the model to play fast rather than to play well.

Be careful reading these percentages: each game is counted against the first criterion it fails, in the order time control, , termination, length. These are not independent rates and they do not add up.

Encoding, and the data leak I narrowly avoided

$PY  encode --games  \
  --vocab  --outdir data
Vocabulaire : 1971 tokens
Corpus      : 
Passage 1/2 — comptage...
  11,035,777 parties, 811,906,492 tokens
  découpage : 10,925,420 parties d'entraînement, 110,357 de validation
Passage 2/2 — encodage...

=== Encodage terminé en 2.2 min ===
  train.bin : 803,780,147 tokens (1.61 Go)
  val.bin   : 8,126,345 tokens (0.02 Go)

Two files, because the data is split into two piles. 99% is used for training, 1% is set aside and will never be shown to the model while it learns. That second pile is the set: it is the exam. It serves to check that the model has extracted general regularities rather than simply memorised what it has seen. A student who revises from past papers and is then examined on those same past papers has proved nothing.

Why only 1%: because 1% of 800 million tokens is still 8 million, plenty to measure precisely. Beyond that, you would be wasting training data without gaining precision. The usual 80/20 rule comes from contexts where you have a few thousand examples; at this scale, it no longer makes sense.

And that is where a trap waits, one I saw coming just in time.

The encoded corpus is not a tidy collection of games. It is a single long ribbon of tokens, where the 11 million games follow one another end to end, separated only by the start and end markers. A flat 1.61 GB file.

The obvious way to build the validation set is therefore to cut that ribbon at 99% of its length and keep the tail. One line of code, and it is done.

Except that the cut falls in the middle of a game. That game, the model saw the beginning of it during training, and it is then asked to predict the rest while pretending this is new data. It knows the exam topic. So it will look better than it is.

What makes this leak dangerous is that it improves the result. There is no crash, no odd curve, no alarm signal. Just a figure that is slightly too good, which you have no reason to suspect. A bug that breaks something always shows up eventually; a bug that flatters your results can survive an entire project.

The countermeasure is simple: instead of cutting the ribbon, I draw 110,357 entire games at random, that is 1%, and each game is wholly on one side or the other. No overlap possible.

A residual contamination, which I publish instead of hiding

For peace of mind, I checked afterwards that no validation game also turned up in training. I compared the exact move sequences of the 10,925,420 training games against the 110,356 validation games.

parties de validation distinctes                    110 356
doublons internes au jeu de validation                    1
parties d'entraînement identiques à une validation      137
taux de contamination                                 0.124 %

137 training games are rigorously identical to a validation game. These are not splitting errors, since I split by game: they are genuine twins. In a corpus of eleven million games, some short and very common opening sequences mechanically end up being replayed identically, move for move, by two players who have never met.

I publish this figure rather than correcting it, for two reasons. The effect on the validation is negligible, at 0.124%. And above all, re-splitting would change nothing: the cause is not in my method, it is in the data itself. A perfectly clean validation set would require deduplicating the corpus, which would raise other questions, notably whether you are entitled to impoverish the distribution you are asking the model to imitate.

It is the kind of check that earns you nothing when it passes, and costs you dearly when you have not done it.

Before launching an hour of compute, I checked that the output was correct rather than merely plausible: 500 games drawn at random, replayed move by move with python-chess from the starting position.

$PY - <<'EOF'
import json, numpy as np, chess
v = json.load(open('')); itos, bos, eos = v['itos'], v['bos_id'], v['eos_id']
d = np.memmap('', dtype=np.uint16, mode='r')
assert d.max() < v['vocab_size'] and d[0] == bos
starts = np.nonzero(d[:200_000] == bos)[0]
bad = games = 0
for a, b in zip(starts, starts[1:]):
    toks = d[a+1:b]
    if len(toks) == 0: continue
    assert toks[-1] == eos
    board = chess.Board()
    for t in toks[:-1]:
        m = chess.Move.from_uci(itos[t])
        if m not in board.legal_moves: bad += 1; break
        board.push(m)
    games += 1
    if games >= 200: break
print(f"{games} parties rejouées, {bad} avec un coup illégal")
EOF
200 parties rejouées, 0 avec un coup illégal

Three minutes of checking that spare you from discovering, a night later, that you trained on corrupted data.

The model: the choices, and why

51,397,120 excluding . Sixteen layers, dimension 512, eight attention heads, a context of 256 moves. For scale, that is roughly a third of the 2019 GPT-2 "small".

Narrow and deep, rather than wide and short

For a fixed parameter count, you can go wide and short, or narrow and deep. I took 16 layers of 512 rather than 7 layers of 768.

The reason is specific to the domain: each additional layer is one more opportunity to recombine information, and a chess move depends on patterns that build up in stages. "This knight is pinned, therefore this piece is free, therefore this square is weak." That is chained reasoning, and it needs depth.

The four modern building blocks

The architecture is not that of the original GPT-2, but that of recent models. Four changes, and each has a real reason to be there.

. You normalise before each sub-layer rather than after. The difference looks cosmetic; it is not: with pre-norm, the residual path that runs through the whole network meets no , so the can travel back from the output to the first layer without being altered. That is what makes deep networks trainable without acrobatics.

RMSNorm. Classic LayerNorm re-centres the values then rescales them. RMSNorm only does the second operation. The re-centring brings almost nothing and costs another pass over the data.

. The question "how does a know what order the moves are in?" is anything but obvious: the attention mechanism treats its inputs as a bag, with no order. The historical approach added a position vector to each . RoPE does something else: it rotates the vectors by an angle proportional to their position. An elegant consequence: when you compute the dot product between a query at position m and a key at position n, the result depends only on the difference m − n. Attention becomes sensitive to relative distances. For chess, that is exactly what you want: a tactical pattern means the same thing at move 10 as at move 40.

. Instead of applying the same transformation to everything that passes through, the network computes two projections and multiplies them term by term, one acting as a tap on the other. So it learns to decide, dimension by dimension and depending on the input, what deserves to be passed on.

The number that stopped me dead

While testing the forward pass, I get a of 5.97 on a randomly initialised model.

That makes no sense. An untrained model should give ln(1971) = 7.59, the loss of a uniform distribution over 1971 , in other words that of someone answering at random. Getting 5.97 would mean knowing something before having seen a single piece of data. That is impossible.

The error was not in the model, it was in my test. I had written model(x, x): I was asking it to predict the current token instead of the next one. Now the are tied, the matrix that turns a token into a vector at the input is the same one that turns a vector back into a token at the output. Combined with the residual connections, which carry the embedding from input to output without major alteration, that matrix recognises on arrival the token it had encoded at departure. So the network knew how to copy its input from initialisation, without learning anything at all.

With the targets correctly shifted, the loss climbs back to 7.71, just above chance. Expected behaviour.

I kept that check in the code, because the lesson is general: an abnormally low loss is never good news, it is a to track down. Here it was benign, a test artefact. In a real training run, the same symptom signals that the answer is reachable from the input, and you notice far too late.

The two-minute test that licenses a two-hour run

This is the safeguard I now consider non-negotiable. Before committing hours of compute, you check that the code can learn a tiny dataset by heart: 100 games, 7,000 .

$PY  --overfit --run-name overfit --overfit-games 100 \
  --overfit-steps 1200 --batch-size 64
step   10 : loss 7.13
step   50 : loss 2.28
step  100 : loss 0.16
step  270 : loss 0.035   -> test réussi, en 2 min 8 s

The starts at chance and collapses. The model has memorised the hundred games, so the training loop, the data pipeline and the are correctly wired.

If it does not fall, it is neither the data nor the : it is a bug. A misaligned target, a mask the wrong way round, a that does not flow. All of them invisible on a full run, discovered only after a night of compute and a mediocre result that then takes hours to diagnose.

The analogy: before printing a thousand pages, you print one. If it comes out blank, no point launching the other 999.

Choosing the batch size: a quarter of the time saved on a single parameter

Before launching, one setting remains to be fixed, and it carries real weight: the batch, that is, the number of sequences processed simultaneously at each step.

During the test, with a batch of 64, the topped out at 53.6%. Moving to 192, it rises to 67.4%, a quarter of the compute time saved for a single changed.

The explanation lies in the fixed costs per step. At each iteration, the GPU does a number of things that do not depend on the batch size: updating the 51 million parameters with , launching the CUDA kernels, synchronising. Those costs are paid once per step, whether the batch holds 64 sequences or 192. By tripling the batch, you amortise them over three times as much useful data. It is the trip to the tip: the driving time is the same whether the boot is half empty or full.

It does not scale indefinitely. saturates, and above all, with too large a batch each weight update becomes less informative. MFU measures the efficiency of the compute, not the efficiency of the learning. The two do not always move in the same direction, and that is a common confusion.

Rather than guessing, I measure what fits in memory and at what throughput:

$PY - <<'EOF'
import os; os.environ["CUDA_DEVICE_ORDER"] = "PCI_BUS_ID"
import time, torch
from model import ChessGPT, ModelConfig
cfg = ModelConfig(n_layer=16, n_embd=512, n_head=8, mlp_hidden=1408)
m = ChessGPT(cfg).to("cuda:1"); opt = torch.optim.AdamW(m.parameters(), fused=True)
for bs in (64, 128, 192, 256):
    torch.cuda.reset_peak_memory_stats(1); torch.cuda.empty_cache()
    try:
        x = torch.randint(0, cfg.vocab_size, (bs, 257), device="cuda:1")
        for i in range(3):
            with torch.autocast("cuda", dtype=torch.bfloat16):
                _, loss = m(x[:, :-1], x[:, 1:])
            loss.backward()
            if i == 0: torch.cuda.synchronize(1); t0 = time.perf_counter()
            opt.step(); opt.zero_grad(set_to_none=True)
        torch.cuda.synchronize(1)
        print(f"batch {bs:>3} : pic {torch.cuda.max_memory_allocated(1)/2**30:5.1f} Go | "
              f"{bs*256/((time.perf_counter()-t0)/2):>8,.0f} tok/s")
    except torch.cuda.OutOfMemoryError:
        print(f"batch {bs:>3} : DÉPASSEMENT"); torch.cuda.empty_cache()
EOF

The rule I apply: take the largest batch that leaves at least 3 GB of headroom. On a run of several hours, that headroom costs nothing in throughput and avoids an abrupt stop for running out of memory an hour and a half in. I settled on 192, which fits in 18.5 GB out of the 24 available.

One last setting before launching, and it is crucial for what follows: --snapshot-every 1000 freezes a of the model every 1000 steps. Without it, there is no way to reconstruct after the fact in what order the model learns. That curve is the most interesting result of the project, and it only exists if you think of it before launching.

The run

nohup $PY  --run-name run1 --device cuda:1 \
  --batch-size 192 --max-steps 16300 --warmup-steps 500 \
  --snapshot-every 1000 \
  > logs/train_run1.log 2>&1 &

The log sets the tone from the very first lines, and it is a moment I rather enjoy: you watch the tumble live.

[données] train 803,780,147 tokens | val 8,126,345 tokens
[modèle] 51,397,120 paramètres non-embedding | 333,548,544 FLOP/token
[optim] AdamW fused | 81 tenseurs avec weight decay, 33 sans
[compile] torch.compile en cours (première itération lente)...
[plan] 49,152 tokens/step x 16,300 steps = 801 M tokens
step     10 | loss 7.4481 | lr 1.20e-05 |    16,533 tok/s | MFU  8.8% |   0.5 M tokens |   0.5 min
step     20 | loss 7.1423 | lr 2.40e-05 |   127,005 tok/s | MFU 67.4% |   1.0 M tokens |   0.6 min
step     50 | loss 6.3802 | lr 6.00e-05 |   127,351 tok/s | MFU 67.6% |   2.5 M tokens |   0.8 min
step    100 | loss 5.5064 | lr 1.20e-04 |   126,370 tok/s | MFU 67.1% |   4.9 M tokens |   1.1 min

The first step is slow (16,533 /s, 8.8%) because torch.compile is busy optimising the model. From the second onwards, we are in steady state: 127,000 tokens/s, MFU 67.4%, 18.5 GB of out of 24, 279 W, 97% GPU utilisation.

A reading taken after 99 minutes of continuous load incidentally validates the 280 W decision. The 30-second benchmark measured 54 °C; the real training run stabilises at 78 °C. Twenty-four degrees of difference, without the slightest throttling: throughput stayed at 120,000 tokens/s from start to finish. A short benchmark tells you nothing about thermal behaviour in steady state, and that is precisely the margin I had kept by not going up to 420 W.

And two hours later, the end:

step  16300 | loss 1.6024 | lr 6.00e-05 |   120,350 tok/s | MFU 63.9% | 801.2 M tokens | 119.8 min
  >> eval step 16300: loss_val 1.6437 (train 1.6024)

=== Terminé : 16300 steps, 801.2 M tokens, 119.9 min ===

In parallel, on the second card, an evaluator grades each as it appears:

nohup $PY  --run-name run1 --device cuda:0 --prefix eval \
  > logs/eval_watcher.log 2>&1 &
Surveillance de checkpoints/run1_step*.pt sur cuda:0
[15:17:19] évaluation du step 1000...
  85.10% [IC95 83.47% - 86.59%] sur 2,000 positions
  terminé en 54 s
[15:24:15] évaluation du step 2000...
  91.55% [IC95 90.25% - 92.69%] sur 2,000 positions

Each evaluation runs in a subprocess rather than by calling the function directly. This is not fussiness: PyTorch never fully returns the GPU memory of a released model within a single process, because of allocator fragmentation. Over twenty or so successive evaluations, you would end up out of memory on a 12 GB card. A process that dies, on the other hand, returns everything.

Testing the code before you need it

The run was going to produce its first snapshot within minutes, and would immediately call , a script that had been written but never run. Rather than discover a crash at that moment, I ran them all on the checkpoint from the test, which has seen only a hundred games and therefore plays nonsense. It does not matter: what is being tested is the code, not the model.

The results on that toy model are instructive:

taux de coups légaux en génération libre     23.0 %
accord top-1 avec l'humain                    7.7 %
roque : coup le plus probable légal          55.0 %
promotion : coup le plus probable légal       0.0 %

Zero percent on promotion: a model trained on a hundred games has quite simply never seen a promotion.

And a contrast I find fundamental: that same model, unable to propose a legal move three times out of four, played 300 in a row without a single illegality as soon as the mask is switched on. You set the of impossible moves to minus infinity before taking the maximum. The model keeps its preferences, but it can no longer play a forbidden move.

That contrast between 23% in and 100% with a mask is exactly what the project sets out to document. The mask produces a player that respects the rules; it says nothing about what the model has understood. That is why the central figure of this article is the one without a mask.

First measurement: 85% after twenty minutes

The first lands at step 1000, that is 49.2 million seen, 6% of the planned training.

85.10% legal moves in , 95% CI [83.47 – 86.59], over 2000 positions.

Twenty minutes of compute, and five times out of six what it proposes is playable. Without ever having seen a chessboard.

The breakdown of the errors says more than the overall rate:

déplacement illégal    209  (coup géométriquement correct, illégal ici)
case de départ vide     54
token spécial           27  (<pad>, <bos> ou <eos> proposé comme un coup)
pièce adverse            8

The dominant pattern is the most interesting one. "Illegal displacement" means the model proposed a move that looks like a real move, the right piece on the right starting square, a displacement consistent with how that piece moves, but impossible in that precise position: destination square occupied, piece pinned, king that would remain in check.

In other words, at this stage the model has already largely understood how the pieces move. What it lacks is reasoning about the current state. It knows the grammar, not yet the context. And only 8 attempts out of 2000 to move an opponent's piece: it already knows whose turn it is.

The result that surprised me most

The per-rule tests, at the same moment:

règle           coup le plus probable légal    masse de probabilité légale
roque                     98.67 %                      96.20 %
prise en passant          98.00 %                      90.19 %
promotion                 88.00 %                      71.63 %
sortie d'échec            76.67 %                      61.93 %

I expected the opposite. Castling and en passant are the two rules explained last to beginners: they are exceptions, with convoluted conditions (the king has not moved, no square crossed is attacked, the opposing pawn has just made a double step on the previous turn). Getting out of check, on the other hand, is the most fundamental constraint of the game, the one without which there is no chess game at all.

And it is the fundamental one the model fails most often.

The explanation lies in the nature of what is learned. Castling is a surface pattern: in , it is always e1g1, e1c1, e8g8 or e8c8, four strings that come back millions of times, almost always at the same point in the game, after the same handful of developing moves. It is a sequence to memorise, and a language model excels at that.

Getting out of check is not. There is no template sequence. You have to know which square is attacked, by what, and which pieces can interpose or capture, which means holding a representation of the position and reasoning about it. No surface pattern can stand in for that.

Here is the full trajectory, between 49 and 786 million :

RuleNature for a networkStartEnd
Castlingfixed pattern, 4 set strings98.7%100%
En passantrare but distinctive pattern98.0%100%
Promotionpattern + one condition88.0%99.3%
Getting out of checkno pattern, state to reconstruct76.7%96.0%

Three rules are born at the ceiling: they have no trajectory, they are acquired from the first steps. Only one truly learns, and it takes around 400 million tokens to catch up with the others. The hierarchy observed at 49 million tokens never changed order across eight hundred million.

The deliberately cruel number

So far I have been measuring isolated moves. Let us put the question differently: across a whole game, how many make it to the end without a single mistake?

1.67%. One in sixty.

The gap with the earlier 85% may be surprising, but the two figures do not answer the same question, and the second is far more demanding. Completing a whole game means getting every move right in a row: a single mistake, anywhere, and the game is disqualified.

And probabilities multiply. Getting two moves right in a row at 85% is 0.85 × 0.85, that is 72%. Three moves, 61%. Ten moves, 20%. And a human game has seventy of them. So the probability collapses well before the end.

It is the classic trap of any system that chains steps together: a respectable success rate per step becomes catastrophic over a long sequence. The same reasoning applies to an agent chaining tool calls, or to a processing chain with ten links.

One clarification, otherwise the arithmetic does not add up. Multiplying 0.85 seventy times would give a far more catastrophic result than the 1.67% measured. Two reasons for the gap, and they are instructive. The model does not make mistakes at random: it is markedly more reliable in the opening, on positions seen millions of times, than in an unfamiliar middlegame. And above all, a generated game also stops when the model itself decides it is over, which it often does too early. Those short games count as successes.

I want to publish both figures. The per-move rate measures what the model has understood; the per-game rate measures its consistency. A model usable in real play needs both, and the second is merciless.

End of the run, and the final evaluation

Training finished: 16,300 steps, 801,177,600 , exactly two hours. Training 1.602, best 1.643. The two curves never diverged: on a single epoch with 800 million tokens for 51 million , there is not enough repetition to overfit.

$PY  --ckpt  --device cuda:1 \
  --n-legal 20000 --n-agreement 20000 --n-per-rule 500 --n-full-games 500 \
  --out 
[1/4] Taux de coups légaux en génération libre...
      97.86% [IC95 97.65% - 98.05%] sur 20,000 positions

[2/4] Accord avec le coup humain...
      top-1 51.42% | top-5 88.72%

[3/4] Tests par règle...
      roque        n= 500 | top1 légal 100.00% | = humain 77.80% | masse légale 99.50%
      en_passant   n= 500 | top1 légal 100.00% | = humain 76.60% | masse légale 98.80%
      promotion    n= 500 | top1 légal  98.60% | = humain 78.40% | masse légale 96.03%
      echec        n= 500 | top1 légal  96.60% | = humain 69.00% | masse légale 91.55%

[4/4] Parties complètes sans coup illégal...
      60.40% [IC95 56.05% - 64.59%] sur 500 parties

Rapport ->  (86 s)

The result, the heart of the project:

97.86% of proposed moves are legal, in , with no constraint whatsoever. 95% CI [97.65 – 98.05], over 20,000 positions.

The interval is four tenths of a point wide, against more than three points on the intermediate measurements at 2000 positions. That is the mechanical benefit of sample size: precision improves as the square root of the number of observations.

règle              coup le plus probable légal    = coup humain
roque                       100.00 %                 77.80 %
prise en passant            100.00 %                 76.60 %
promotion                    98.60 %                 78.40 %
sortie d'échec               96.60 %                 69.00 %

The number not to over-interpret

60.40% of generated games unfold without a single illegal move, against 1.67% at the first . A spectacular improvement, and the easiest figure in the project to misread:

le modèle a prédit <eos>                275 parties   55.0 %
un coup illégal a interrompu la partie  198 parties   39.6 %
vraie fin de partie selon les règles      27 parties    5.4 %

The vast majority of the "successes" are games the model declared finished by emitting the end token, not games carried through to mate. So I publish both figures with their breakdown, rather than the bare 60.40% which would suggest the model takes six games out of ten all the way to mate.

Agreement with humans, and why it is not a measure of strength

51.42% at , 88.72% at top-5. One time in two, the model proposes exactly the move a human between 1800 and 2600 played.

An objection comes to mind immediately, and it is a healthy one: since the model was trained on human moves, is it not normal that it recovers human moves? Where is the merit?

The answer comes in two parts.

It has never seen these positions. The measurement is taken on the validation set, that famous 1% set aside at the start. These are games that were never shown to the model during training. So it is not reciting, it is generalising: from an unfamiliar sequence of moves, it has to guess what a human would play next. For scale, choosing at random among the legal moves would give roughly 3% success; it is at 51%.

And that is very precisely its task. Predicting a human's next move is the objective it was trained on, no more and no less. So this figure measures how well it does the job it was asked to do. It is the equivalent, for a language model, of "guess the next word in this sentence".

On the other hand, you have to resist the temptation to turn it into a measure of playing strength. That is the most tempting misreading, and it is spectacularly wrong. A perfect engine like would have low agreement with 2000 Elo players, precisely because it plays better than they do. It would often choose a move the human did not find, and would therefore be counted as disagreeing.

In other words, this figure rises when the model resembles the humans in its corpus, not when it plays well. Imitating and excelling are two distinct objectives, and it happens that my model pursues only the first. That is also why the real measure of strength will come from elsewhere: from games won or lost, against Stockfish and then against real players.

The mistake that cost me most, and that saved fifteen hours of compute

Here is the one I am least proud of, and the most instructive.

I had measured that the model generated games of 46.5 against 71.6 in the corpus, and that only 5.4% reached a terminal position. I had concluded that it did not know how to play endgames, and I was about to launch a second training run with endgames oversampled. Fifteen hours of compute.

Before launching, I checked the diagnosis. It was wrong, on two distinct counts.

First bias: I had measured at the wrong . The evaluation was running at temperature 1.0 while the engine itself plays at 0. At temperature 1, the model puts roughly 0.9% of probability on <eos> at each position. Over seventy half-moves, that small percent accumulates, and the probability of hitting <eos> by accident before the end becomes the majority case.

A temperature sweep, 200 games per point:

température   parties sans faute   longueur   fins réelles / 200
    1.0             52.5 %           48.2            9
    0.8             68.0 %           54.5           19
    0.6             80.5 %           64.1           24
    0.4             82.5 %           63.6           33
    0.2             87.0 %           60.5           39

At temperature 0.6, games run to 64.1 half-moves against 71.6 in the corpus, that is 89% of the real length and not the 65% I had announced. The rate of games without a mistake goes from 52.5% to 87%.

Second bias, more serious: I was comparing against the wrong reference. I had implicitly compared those 5.4% against 100%. Except I had never measured what humans do. Replaying 3000 real games from the set:

finissant sur une position terminale     29.0 %   (dont mat : 23.2 %)
finissant par abandon ou accord de nulle 71.0 %

Seven human games out of ten end in a resignation, not in mate. And in my stream of , a resignation is written exactly <eos>: there is no terminal position to reach. When the model emits <eos> without mate, it is not shirking, it is reproducing the dominant behaviour of its data.

At temperature 0.6 it produces 12% terminal endings against 29% for humans. So it resigns a little too early, but by a factor of two, not a factor of twenty. The idea was dropped: it would have consumed fifteen hours to fix a largely non-existent problem, and would probably have degraded the model by distorting the distribution it is asked to imitate.

Three conditions made this mistake nearly undetectable: the setting looked secondary, the result pointed the way I was hoping, and the gap looked enormous and therefore significant. Temperature is not a of the model, it is a parameter of the measurement. Since then, a temperature fixed per metric:

MetricTemperature imposed
Legality per move1.0 (harshest test)
, per-rule testsnot applicable (computed on the )
Complete games1.0 and 0.6, both published
Matches and 0 (deterministic)

What it extracted without being asked

Once the model was trained, I gave it e2e4 c7c5, the Sicilian. And I looked not at its move, but at its whole distribution, renormalised over the legal moves:

ReplyProbability
Nf3, the Open Sicilian54.7%
Nc3, the Closed variation10.8%
c3, the Alapin8.9%
d4, the central push7.4%
Bc4, the Italian bishop5.8%
f4, the Grand Prix Attack4.2%

That is, in exactly that order, the popularity ranking of these six replies among human players. The move it rates most probable is also the most played in the world; the last one, the rarest.

Nobody gave it that ranking. It knows neither the names of the variations nor their reputation. It has only seen, millions of times, what people play. That is what learning by prediction is: the model compresses the collective preferences of its data into a distribution.

That ranking reflects players' taste, not the truth of the position. The model imitates popularity, it does not compute the best move. The distinction holds for any system trained by imitation.

Its real strength

brew install stockfish        # ou apt, ou compilation

$PY  --ckpt  --device cuda:1 \
  --stockfish $(which stockfish) --levels 0 1 2 3 --games 200 --movetime-ms 50 \
  --out 
  == Stockfish Skill Level 0 ==
    -> score 56.0% (+97 =30 -73) | Elo +42 ± 45
  == Stockfish Skill Level 1 ==
    -> score 37.8% (+56 =39 -105) | Elo -87 ± 44
  == Stockfish Skill Level 2 ==
    -> score 26.2% (+33 =39 -128) | Elo -180 ± 48
  == Stockfish Skill Level 3 ==
    -> score 18.0% (+17 =38 -145) | Elo -263 ± 52

800 games, alternating colours and four opening drawn at random so that the games are not all identical:

OpponentScore gap
level 056.0%+42 ± 45
Stockfish level 137.8%−87 ± 44
Stockfish level 226.2%−180 ± 48
Stockfish level 318.0%−263 ± 52

The +42 ± 45 at level 0 is a statistical tie, the interval contains zero.

And by having each play against the final model, you get the progression across the whole run:

$PY  --ckpt  --device cuda:1 \
  --skip-stockfish --ladder --ladder-run run1 --ladder-games 60 \
  --out 
step   1000 |   49.2 M tokens | Elo  -636.4          (aucune victoire sur 60)
step   2000 |   98.3 M tokens | Elo  -338.0 ± 102
step   4000 |  196.6 M tokens | Elo  -279.6 ± 106
step   8000 |  393.2 M tokens | Elo  -127.0 ±  77
step  12000 |  589.8 M tokens | Elo   -94.9 ±  79
step  16000 |  786.4 M tokens | Elo   +52.5 ±  76
23 min, 60 parties par instantané

The principle is simple: each snapshot faces the final model. That gives an internal scale where the final model serves as the zero mark. The very first snapshot sits 636 points below, the best one ends 52 points above, that is 688 points of progression in two hours of compute.

Why I do not say "my model is 1500 Elo"

It is worth pausing on what these numbers mean, because they are the kind that gets quoted wrongly.

Elo never measures an absolute strength, only a gap between two players. It is a relative scale, like saying "I am 20 centimetres taller than my brother". The information is exact, but it only gives my height if you already know my brother's.

Here, my two points of comparison are a throttled Stockfish and my own snapshots. So I can state without reservation that my model gained 688 points on itself, and that it holds its own against Stockfish at its weakest level. On the other hand, I cannot derive a Lichess or FIDE rating from it, because I have measured no one whose height I already know. That reference point is what is called an .

It would have taken, for instance, having my model play an opponent whose Lichess rating is established, and I could then have placed mine on that scale. I did not do it at this stage, so my results file states in plain words "absolute_elo": "not measured" rather than an estimate that would look impressive while resting on nothing.

That is also why the rest of the project is interesting: by putting the model online on Lichess, it finally gets an anchored rating, computed against thousands of real opponents. I no longer need to estimate, I just need to read.

A second model, and a question that caught me out

I wanted a second run, a bigger one. I had first sized it on three months of data, then I asked myself why not four, five or six.

The question contained a confusion worth clearing up. You do not "improve" the optimum by adding data at a fixed model size: you go past it. Chinchilla does not say "twenty per is the best possible model", it says "for a given compute budget, split it this way between size and data". Training beyond it gives a better model, simply a less efficient one per FLOP spent. That is exactly what Llama and its peers do, training eight-billion-parameter models on fifteen trillion tokens, a hundred times the Chinchilla optimum. They do it because the cost that matters to them is not training, paid once, but , paid on every request.

In costing the options, what I mostly discovered was that my initial constraint was a tacit assumption, never checked: I had told myself "I'll launch this tonight", and had sized everything on one night without ever asking the question.

modèle fixe à 120 M          modèle redimensionné à chaque corpus
 3 mois : 14.7 h              3 mois : 120 M -> 14.7 h
 4 mois : 19.9 h              4 mois : 142 M -> 23.0 h
 5 mois : 25.1 h              5 mois : 214 M -> 41.3 h
 6 mois : 30.3 h              6 mois : 252 M -> 57.8 h

The run started again on four months of data and 141,589,248 parameters, 20 layers, dimension 768, 12 heads, that is 21.6 tokens per parameter.

The full pipeline fits in , but here is what it actually executes, step by step. It is the same sequence as for the first model, with three additional months of data.

# 1. Parser les trois mois supplémentaires (juillet est déjà fait)
for m in 2026-06 2026-05 2026-04; do
  $PY  parse \
    --input data/lichess_db_standard_rated_${m}.pgn.zst \
    --workers 10 --target-tokens 0 --max-games 0 \
    --out data/games_uci_${m}.txt --stats logs/phase1_parse_${m}.json
done

# 2. Fusionner les quatre mois en un seul corpus
cat  data/games_uci_2026-06.txt \
    data/games_uci_2026-05.txt data/games_uci_2026-04.txt \
    > data/games_uci_4mois.txt

# 3. Encoder ce corpus dans son propre répertoire
$PY  encode --games data/games_uci_4mois.txt \
  --vocab  --outdir data/4mois

# 4. Entraîner le modèle de 142 M, une époque sur les 3,19 G tokens
$PY  --run-name run2 --device cuda:1 --data-dir data/4mois \
  --n-layer 20 --n-embd 768 --n-head 12 --mlp-hidden 2048 \
  --batch-size 80 --grad-accum 2 \
  --max-steps 77787 --warmup-steps 2000 \
  --eval-interval 1000 --snapshot-every 5000
[données] train 3,186,155,520 tokens | val 32,183,552 tokens
[modèle] 141,589,248 paramètres non-embedding | 896,721,408 FLOP/token
[plan] 40,960 tokens/step x 77,787 steps = 3,186 M tokens

Going from 17 to 20 layers changes the memory footprint, and I did not want to assume it. Measured: at batch 96 the peak reaches 22.8 GB out of 24, at batch 80 it falls to 19.3 GB, for a rigorously identical throughput (34,598 against 34,884 tokens/s). The safety margin was free. On an eighteen-hour run, three gigabytes of cushion against zero percent of performance, there is nothing to hesitate about.

What this second model gives

=== Terminé : 77787 steps, 3186.2 M tokens, 1078.7 min ===
résumé -> logs/run2_train_summary.json

Final training 1.4312, best loss 1.4439.

Eighteen hours of training, 3.19 billion tokens seen, that is four times more than the first model. It is then evaluated exactly like the first:

$PY  --ckpt  --device cuda:1 \
  --vocab  --val-games data/4mois/val_games.txt \
  --n-legal 20000 --n-agreement 20000 --n-per-rule 500 --n-full-games 500 \
  --out logs/eval_final_run2.json
coups légaux en génération libre   98.85 %                       n=20 000
accord humain top-1 / top-5        55.49 % / 91.95 %             n=20 000

roque             100.00 %      prise en passant  100.00 %
promotion          99.80 %      sortie d'échec     98.60 %

The comparison, under the same measurement protocol ( 1.0 on both sides):

run1run2
Parameters51 M142 M
Data seen801 M tokens3,186 M tokens
Duration2 h18 h
Training loss1.6021.431
Validation loss1.6431.444
Legal moves97.86%98.85%
51.42%55.49%
Human agreement top-588.72%91.95%

The second model is better everywhere, that much is clear. But nine times the compute for one point of legality is a return that raises questions, and the gap between the two right-hand columns is the real subject. That is what we are going to look at now, and the result caught me off guard.

The most interesting result of the project

Under a strictly equal protocol, and above all at the same number of seen, which neutralises the effect of data volume and leaves only that of size:

M tokens |  run1 légaux  run1 top-1 |  run2 légaux  run2 top-1
         |   (51 M param.)          |   (142 M param.)
     205 |     95.10 %      46.55 % |     95.05 %      49.20 %
     410 |     96.90 %      49.10 % |     97.10 %      50.30 %
     614 |     97.55 %      50.95 % |     97.35 %      52.75 %
     819 |     97.75 %      51.85 % |     97.40 %      53.80 %

Two columns tell two opposite stories.

On legality, tripling the model size changes nothing. The gaps are within the noise, sometimes in favour of the small model.

On agreement with the human move, the gap is clear and constant: two to three points in favour of the big model, at every level, without exception.

The interpretation that seems solid to me: learning the rules is an easy task, which a small model saturates. The rules of chess are a finite, regular system with no exceptions. Once those regularities are extracted, which takes around 250 million tokens, there is nothing left to learn, and the extra capacity is of no use. The ceiling is not in the model, it is in the task.

Predicting which move a human would play is of another nature. It is no longer mechanics, it is judgment: evaluating a position, arbitrating between several reasonable plans, reproducing the preferences of a population. That task has no visible ceiling, and it absorbs all the capacity you give it.

Put differently: size does not buy knowledge of the rules, it buys discernment.

And what if we make them play each other?

The metrics say the big model is slightly better. It remained to be seen what that gives on the board. So I had them face each other directly, 400 games, alternating colours, four opening drawn at random for variety:

$PY  --ckpt  --device cuda:1 \
  --skip-stockfish --vs  --vs-games 400 \
  --label-a "run2 (142 M, 4 mois)" --label-b "run1 (51 M, 1 mois)" \
  --out logs/duel_run2_vs_run1.json
  == Duel direct : run2 (142 M, 4 mois) contre run1 (51 M, 1 mois), 400 parties ==
    -> run2 marque 84.1% (+305 =63 -32) | Elo +290 ± 40

  terminaisons : 337 mats, 47 nulles par répétition,
                 13 pats, 3 par matériel insuffisant

305 wins against 32, that is a 290-point gap. Two to three extra points of therefore translate into a very clear superiority in play. What read as a modest gap on a metric becomes an unambiguous balance of power on the board.

And a detail worth noting: across those 400 games, not one ended on an illegal move. All of them reached a genuine end of game, 337 of them by checkmate. These are two models that have never seen a chessboard, playing each other for four hundred games, and going to mate without ever breaking a rule nobody gave them.

And here is the detail that sends a chill down your spine. This result would have been invisible without the correction described earlier. With the run2 watcher launched by mistake at 0.6, run2 showed 99.20% then 99.35% legality against 97.86% for run1. I would have concluded that scale brought a spectacular gain on legality, that is, exactly the opposite of what the data shows under an equal protocol. A measurement setting changed between two experiments would have produced a false conclusion, a spectacular one, pointing in the direction I was hoping for. The three conditions met for it never to be questioned.

The model is a Llama without knowing it

While trying to publish the model on Hugging Face, I discovered there was almost nothing to convert. + + + is the Llama architecture. So the model reloads as a standard LlamaForCausalLM, not through tinkering, but through structural identity.

$PY -m pip install transformers safetensors
$PY  --ckpt  --out hf/chess-51m
  longueur   1 : écart max sur les logits = 0.00e+00
  longueur  16 : écart max sur les logits = 0.00e+00
  longueur  64 : écart max sur les logits = 0.00e+00
  longueur 256 : écart max sur les logits = 0.00e+00

[OK] écart maximal 0.00e+00 — les deux implémentations sont identiques

Écrit dans hf/chess-51m/ (52.4 M paramètres)

The conversion is verified numerically: the script refuses to write if the of the two implementations differ by more than 1e-4. Measured: 0.00e+00 at every sequence length, that is a bit-for-bit match and not a mere approximation.

The unusual part is the tokenizer, since there is no but a table of 1971 moves. That is also what causes trouble with llama.cpp, whose use I advise against here for that reason.

142 million against 35 billion

I had the model play against Qwen3.6-35B, served locally by llama-server on its OpenAI-compatible API, under identical conditions.

sudo systemctl start llama.service
curl -s http://localhost:8080/health        # attendre {"status":"ok"}

$PY  --ckpt  --device cuda:1 \
  --games 100 --retries 3 --out 

And to make the comparison as unfavourable as possible to my model, I grant Qwen a considerable handicap, one it would have in no real game:

  • up to 30 attempts for each move, where my model gets a single pass;
  • after each failure, it is told that its move is illegal and reminded of the ones it has already proposed in vain;
  • past ten tries, it is explicitly asked to give any legal move at all, even at random.

Here is what that favourable treatment gives:

coups demandés                     2 275
obtenus au premier essai             823   (36.18 %)
finalement obtenus                 1 688
JAMAIS obtenus, même en 30 essais    587

tentatives totales                24 868   (10.9 par coup en moyenne)
propositions impossibles          23 180

These three numbers are worth pausing on. Across 2,275 positions, 587 moves could never be obtained, despite thirty tries, explicit error feedback and permission to answer at random. And to extract the other 1,688, it took 23,180 impossible proposals. A legal move cost eleven attempts on average.

And the game results, handicap included

This is where the exercise makes full sense. Thanks to that handicap, the games no longer stop on a forfeit: they go to the end. So we can finally look at the only verdict that counts in chess, that of the board.

philidor-142M  contre  Qwen3.6-35B, 100 parties
  99 victoires, 1 nulle, 0 défaite
  terminaisons : 99 par échec et mat, 1 par pat
  écart d'Elo : +919  (borne basse de l'intervalle à 95 % : +730)

philidor-51M   contre  Qwen3.6-35B, 100 parties
  98 victoires, 2 nulles, 0 défaite
  terminaisons : 98 par échec et mat, 1 pat, 1 nulle par répétition
  écart d'Elo : +798  (borne basse : +646)

Zero forfeits. Zero losses. The wins are mates, on the board. And the small 51-million- model, the two-hour one, does just as well as its big brother: 98 mates against 99.

So here is the full picture. Same question put to all three models, none has the list of legal moves, and the legality rate counts only the first answer:

ModelSizeNatureLegal movesScore over 100 games
Philidor51 Mspecialised97.86%98 w / 2 d / 0 l
Philidor142 Mspecialised98.85%99 w / 1 d / 0 l
Qwen3.635 Bgeneral-purpose36.18%0 wins

My smallest model is 681 times lighter than Qwen. It produces nearly three times as many legal moves, and it mates it 98 times out of 100 even though the opponent has thirty attempts per move, error feedback, and permission to play at random.

It is the most counter-intuitive result of the project, and it fits in one sentence: on a narrow, well-defined task, specialisation crushes size. Qwen knows thousands of things my model has no idea about, it speaks twenty languages and writes code. But it has never had to hold a chessboard in its head, whereas mine did nothing else for two hours, on a gaming card.

One nuance all the same, and it matters if you are not to over-interpret. If Qwen loses, it is not because it plays badly in the strategic sense: it is that it does not know where the pieces are. Two legal moves out of three escape it on the first try, and it proposes empty starting squares or opposing pieces.

One detail confirms this clearly: out of the 2,275 answers, zero are unreadable. Qwen understands the instruction perfectly, always answers in the right format, and produces a perfectly well-formed move. It is not mistaken about the task, it is mistaken about the move. That is the difference between a model that has not understood what is being asked and a model that has understood very well but lacks the internal representation needed to answer. Its own reads the game as a sequence of plausible symbols; it has never built a chessboard. A player who cannot see the board cannot win, whatever their general knowledge. So the correct formulation is not "my model is better at chess than a 35-billion-parameter model", but this one: under identical conditions, and even granting the other a massive handicap, the specialised model plays and wins, the general-purpose one cannot manage it.

Two protocol settings, and a surprise

The list of legal moves is not given to the opponent. Giving it would amount to offering it the mask that my model does not have either in .

And Qwen's reasoning is disabled (enable_thinking: false). This point deserves a word, because the result surprised me. Enabled, the model produces around 530 of thinking before each move, that is 7.7 seconds per move, that is nine hours for a hundred games. Disabled, it answers in 0.4 seconds. I compared the two on the same positions:

défaut                  7.7 s | 527 tokens | 'b8c6'
enable_thinking=false   0.4 s |   5 tokens | 'b8c6'

Exactly the same answer, twenty times faster. Five hundred tokens of reasoning to arrive at the move it would have given straight away. That is a useful reminder at a time when "reasoning" is being added everywhere: on a task where the model lacks the necessary representation, thinking for longer compensates for nothing. Qwen was not short of compute time, it was short of knowing where the pieces are.

Going live, or how I got my own bot banned

The ultimate test remained: real humans. I set up a Lichess bot.

The model worked first time. It was all the plumbing around it that broke.

To adjust the configuration, I restarted the bot half a dozen times in a few minutes. Each start opens a connection to the Lichess event stream, /api/stream/event, which is rate-limited. That burst triggered a 429 Too Many Requests that lasted more than an hour.

The symptom was misleading: the bot accepted challenges (that goes through another endpoint) but never received the "the game has started" event, so it did not play. From the outside, it looked crashed. From the log, it was going round in circles on its internal timer, very much alive.

I explored several false leads: several concurrent instances (no, only one), a bot hammering the network (false, measured: 14 real network calls against 811 local waits). I waited 3 minutes, then 15, then 30 of total silence. Nothing helped: a fresh process probes again immediately and picks up another 429.

The fix is not a naive automatic restart, which would have replayed exactly the hammering that got me banned. It is a watchdog that restarts the bot only if it really dies, with 120 seconds of waiting between attempts. On a rate limit, the bot does not die, it waits: the watchdog must on no account get involved.

The moral is uncomfortable: the model, the hard part, worked first time. It was by trying to go fast on the easy part that I broke everything. A remote service has quotas, respecting them is not optional, and nervously restarting a process is exactly the reflex that turns a setting to be adjusted into an hour-long ban.

A single model for fifty games

Before the results, a word on the architecture, because it holds a pleasant surprise.

The first reflex would be to launch one instance of the model per game. That would be a mistake: each copy occupies its own memory, and you saturate the GPU after a few games. So I separated the model from the game. A single server loads the model once, around 2.5 GB on the small card, the RTX 3060. Each game launches a lightweight client that sends it the list of moves and receives the move to play, over a local socket.

client -> serveur : {"moves": ["e2e4", "e7e5", "g1f3"]}
serveur -> client : {"move": "b8c6"}

What makes the whole thing comfortable is the cost of a move. My model computes no variations, it does only a single forward pass: a few milliseconds. No search, no tree to explore, no thinking time. It answers, full stop.

The consequence is counter-intuitive. In a game, a player spends most of their time waiting for the opponent. Even with fifty games open simultaneously, very few are waiting on a move from us at the same instant. So the server receives a small trickle of requests, each served in a few milliseconds, and it is enough to handle them one after another. I set concurrency to 50 simultaneous games without memory flinching, since there is still only one model loaded.

And the detail that amuses me most: over a whole game, the bot consumes a few seconds of compute in total. It could play games of under ten seconds without being any worse than in rapid, because it does not think for longer when given more time. Where a classical engine collapses at fast time controls for lack of time to search, and where a human plays worse under pressure, my model is rigorously indifferent to the clock. Its level in bullet and in rapid should therefore be identical.

The figures say otherwise, and that is precisely what makes what follows interesting.

Its rating as of 14 August 2026

The bot has been running since, accepting challenges, playing bullet, blitz and rapid. Its rating is not manufactured by me: it is public, established over 691 games that anyone can look up, against opponents who do not know they are facing a language model.

Time controlRatingGames
Rapid1719190
Bullet1681196
Blitz1544215

Zero illegal moves across the 691 games, thanks to the mask that the finite makes possible.

An average rating always hides peaks. Its finest win is a rapid game against Flower-Queen, a bot rated 1965, that is 241 points above it. The opponent is nothing like a new account with an unstable rating: it has existed since June 2025, has played more than 7,400 games, and its rapid rating is established over 1,205 of them.

Here it is in full, move by move. My model has White.

philidor-142M (1724) against Flower-Queen (1965), rated rapid game of 12 August 2026. One hundred and seventy-one half-moves, won by checkmate.
philidor-142M (1724) against Flower-Queen (1965), rated rapid game of 12 August 2026. One hundred and seventy-one , won by checkmate.

And the most interesting part is how it ends:

The final position: 86. Qhh1#, with two queens. The promoted pawn is the one delivering mate.
The final position: 86. Qhh1#, with two queens. The promoted pawn is the one delivering mate.

Eighty-six moves, an endgame carried through to the end, a pawn promoted to a queen, and mate with two queens. Remember what I said earlier: the model tends to declare the game over too early, and endgames are where memorisation is no longer enough, for lack of statistical regularity to exploit. This game is exactly the counter-example. It did not shirk, it pushed its pawn, promoted, and mated.

So a model that cannot see the board, that decides in a few milliseconds without exploring a single variation, holds an eighty-move endgame against a markedly higher-rated opponent, without playing a single illegal move.

The game is public and replayable move by move: lichess.org/delDiQ5h. That is the best possible guarantee against a doctored figure, and the same is true of the other 690.

Yet I have just explained that the model is indifferent to the clock: it plays in exactly the same way, in a few milliseconds, whatever the time control. Its three ratings ought therefore to be identical. They spread over 175 points.

The explanation is not in the model, it is across the board from it. It is its opponents who change. In rapid, a human has time to think, and the model faces players close to their best level: it holds at 1719. In blitz, the same players become less accurate, but they make up for it with fast patterns and good clock management, two things my model does not do at all. It cannot play the clock, never provokes an opponent short of time, and does not speed up when it should.

In other words, its rating does not only measure its strength: it measures the relative strength of its opponents in each time control, and its own total absence of clock play. It is a useful reminder that the same model, strictly identical, gets three different marks depending on the conditions of the test.

What I take away from it

A system with a single instruction, guess the next symbol, reconstructs on its own a surprising share of a mechanism that was never described to it: the geometry of the pieces, the exceptions, a hierarchy of preferences, enough to hold the comparison with a throttled engine and to rate 1700 against humans.

It "understands" nothing, in the sense we mean. It extracts regularities. But the order in which it extracts them says more about neural networks than the final score does.

Five lessons I am keeping, none of them specific to chess.

You never plan against a manufacturer's figure. My card was running at 40% of its power and nothing flagged it.

On a data pipeline, intuition about what is expensive is almost always wrong. I was about to parallelise a step that was already twenty-five times too fast.

An abnormally low is never good news. It is a to track down.

A model's size does not buy what you think. It brought nothing on the rules, everything on discernment. Before scaling a model up, you need to know whether the target metric is limited by capacity or by the difficulty of the task.

The hard part is not training a model, it is measuring it honestly. My costliest mistake was not in the model, it was in my thermometer: two different , and I was about to burn fifteen hours fixing a problem that did not exist, while masking the most interesting result of the project.

And finally, the most counter-intuitive thing in the project. This 51-million- model, trained in two hours on a gaming card, mates a general-purpose model 681 times its size 98 times out of 100, one granted thirty attempts per move at that. Not because it is better in general: Qwen knows a thousand things it will never know. Because it learned the right thing, on the right data, for a single task.

That is perhaps the most useful lesson to keep at a time when everything is measured in billions of parameters: size is not a strategy, it is a means. On a narrow, well-posed problem, a small specialised model, trained for a few euros of electricity, can do better than the giant, and on hardware everyone already owns.


Reproducing the experiment, end to end

Every command below was actually run. The durations are measured, not estimated.

The scripts called (, , , , and the others) were written for this project. You will find them in full in the appendix, at the end of this article, each with a copy button. There is nothing else to fetch elsewhere: everything needed to redo the experiment is on this page.

# ---- 0. Environnement -------------------------------------------------
# Prévoir ~150 Go libres pour un mois de données, ~400 Go pour quatre.
mkdir -p ~/chess-model && cd ~/chess-model
mkdir -p data checkpoints logs

python3 -m venv venv
PY=venv/bin/python
$PY -m pip install --upgrade pip
$PY -m pip install torch --index-url https://download.pytorch.org/whl/cu126
$PY -m pip install python-chess zstandard numpy tqdm matplotlib

# Versions utilisées : torch 2.13.0+cu126, python-chess 1.11.2,
# zstandard 0.25.0, numpy 2.2.6, matplotlib 3.10.9, tqdm 4.70.0

# ---- Le piège à vérifier AVANT ----------------------------------------
# Le GPU est-il bridé ?
nvidia-smi --query-gpu=name,power.limit,power.default_limit --format=csv
sudo nvidia-smi -i 1 -pl 280          # ne survit pas au redémarrage

$PY                       # < 1 min, donne les TFLOPS réels

# ---- 1. Vocabulaire ---------------------------------------------------
# Le plus simple : copier  depuis cet article et l'enregistrer
# sous data/vocab.json. Il ne dépend d'aucune donnée, il est identique
# pour tout le monde.
# Sinon, le régénérer donne exactement le même fichier :
$PY            # 1971 tokens, instantané

# ---- 2. Données -------------------------------------------------------
cd data
wget -c https://database.lichess.org/standard/lichess_db_standard_rated_2026-07.pgn.zst
cd ..

$PY  benchmark --games 50000        # 2 min, NE PAS SAUTER

nohup $PY  parse --workers 10 --target-tokens 800e6 \
  --out  --stats logs/phase1_parse_stats.json \
  > logs/parse_full.log 2>&1 &                     # 54 min

$PY  encode --games  \
  --vocab  --outdir data            # 2.2 min

# ---- 3. Le garde-fou (obligatoire) ------------------------------------
$PY  --overfit --run-name overfit --overfit-games 100 \
  --overfit-steps 1200 --batch-size 64             # 7.13 -> 0.035 en 2 min
# Si la loss ne tombe pas sous 0.05 : il y a un bug. Ne pas continuer.

# ---- 4. Entraînement --------------------------------------------------
nohup $PY  --run-name run1 --device cuda:1 \
  --batch-size 192 --max-steps 16300 --warmup-steps 500 \
  --snapshot-every 1000 > logs/train_run1.log 2>&1 &     # 2 h exactement

# Évaluation continue sur la seconde carte (optionnel mais précieux :
# la courbe d'apprentissage est impossible à reconstituer après coup)
nohup $PY  --run-name run1 --device cuda:0 --prefix eval \
  > logs/eval_watcher.log 2>&1 &

# Reprendre après interruption (checkpoints toutes les 30 min)
$PY  --run-name run1 --device cuda:1 --batch-size 192 \
  --max-steps 16300 --warmup-steps 500 --snapshot-every 1000 --resume

# ---- 5. Évaluation ----------------------------------------------------
$PY  --ckpt  --device cuda:1 \
  --n-legal 20000 --n-agreement 20000 --n-per-rule 500 --n-full-games 500 \
  --out                        # 86 s

# ---- 6. Force de jeu --------------------------------------------------
$PY  --ckpt  --device cuda:1 \
  --stockfish $(which stockfish) --levels 0 1 2 3 --games 200 \
  --movetime-ms 50 --out       # ~36 min

$PY  --ckpt  --device cuda:1 \
  --skip-stockfish --ladder --ladder-run run1 --ladder-games 60 \
  --out                        # 23 min, 688 points d'Elo

# Duel entre deux modèles
$PY  --ckpt  --device cuda:1 \
  --skip-stockfish --vs  --vs-games 400 \
  --label-a "run2" --label-b "run1" --out logs/duel.json

# Duel contre un LLM généraliste. Le protocole par défaut (dernière ligne)
# accorde 30 tentatives par coup à l'adversaire : c'est celui de l'article.
# Les deux premiers sont plus sévères, les parties s'arrêtent plus tôt.
$PY  --strict --games 100 --out logs/duel_strict.json
$PY  --max-attempts 4 --games 100 --out logs/duel_4.json
$PY  --games 100 --out 

# ---- 7. Graphiques ----------------------------------------------------
$PY  --run-name run1
$PY  --run-name run1 --comparer run1 run2
$PY 

# ---- 8. Utiliser le modèle --------------------------------------------
$PY  --ckpt  --device cuda:0
$PY  --noir --analyse                       # jouer noir, avis affiché

$PY  --ckpt  --selftest
printf 'uci\nisready\nposition startpos moves e2e4\ngo\nquit\n' | \
  $PY  --ckpt 

$PY -m pip install transformers safetensors
$PY  --ckpt  --out hf/chess-51m

# ---- 9. Le second run, plus gros --------------------------------------
nohup  > logs/run2_pipeline.log 2>&1 &
nohup  > logs/post_run2_plots.log 2>&1 &

The pitfalls, summarised

PitfallSymptomFix
Power-capped GPUthroughput 2 to 3× below the spec sheetnvidia-smi -pl
Pool.imap on a OOM killer after a few minutesprocess in fixed-size windows
Train/val split by too good, no symptomsplit by game
Abnormally low loss5.97 instead of 7.59 at initialisationlook for a
Throughput measured on the first stepsduration estimate optimistic by ~15%measure in steady state
Different between two measurementsartificial gap attributed to the modelone temperature per metric
Rapid restarts of a bot429 lasting more than an hourslow watchdog, on real crashes only

All the project's scripts

The 24 files written for this project, in the order they are used. Each one folds open and copies with a single click. Everywhere in the article, an underlined file name opens its code directly in a window, without losing your place in the reading.

bench_gpu.py125 lines

Checks CUDA and measures the real bf16 TFLOPS of each card. This is the one that revealed the 170 W power cap.

"""Phase 0 — Vérification de l'environnement et benchmark matmul bf16.

Mesure les TFLOPS bf16 réels de chaque GPU visible, afin de pouvoir estimer
plus tard la durée de l'entraînement (via le MFU : Model FLOPs Utilization).

On force CUDA_DEVICE_ORDER=PCI_BUS_ID pour que les index PyTorch correspondent
exactement à ceux affichés par nvidia-smi (sinon PyTorch trie par "GPU le plus
rapide d'abord", ce qui prête à confusion quand on a deux cartes différentes).
"""

import json
import os
import platform
import subprocess
import sys
import time
from datetime import datetime, timezone

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import torch


def nvidia_smi_table():
    out = subprocess.run(
        ["nvidia-smi",
         "--query-gpu=index,name,memory.total,memory.used,memory.free",
         "--format=csv,noheader,nounits"],
        capture_output=True, text=True, check=True).stdout.strip()
    rows = []
    for line in out.splitlines():
        idx, name, total, used, free = [p.strip() for p in line.split(",")]
        rows.append({"index": int(idx), "name": name,
                     "memory_total_mib": int(total),
                     "memory_used_mib": int(used),
                     "memory_free_mib": int(free)})
    return rows


def bench_matmul(device_idx, n=8192, dtype=torch.bfloat16, warmup=5, iters=30):
    """Multiplie deux matrices n x n et renvoie le débit en TFLOPS.

    Un produit matriciel n x n coûte 2*n^3 opérations flottantes (n^3
    multiplications + n^3 additions). En mesurant le temps, on en déduit le
    nombre d'opérations par seconde que la carte soutient réellement --
    toujours bien en dessous du chiffre marketing du constructeur.
    """
    dev = torch.device(f"cuda:{device_idx}")
    free_bytes = torch.cuda.mem_get_info(dev)[0]
    # 3 matrices (a, b, résultat) de n*n éléments sur 2 octets, + marge x2
    need = 3 * n * n * 2 * 2
    while need > free_bytes and n > 1024:
        n //= 2
        need = 3 * n * n * 2 * 2

    a = torch.randn(n, n, device=dev, dtype=dtype)
    b = torch.randn(n, n, device=dev, dtype=dtype)

    for _ in range(warmup):
        c = a @ b
    torch.cuda.synchronize(dev)

    t0 = time.perf_counter()
    for _ in range(iters):
        c = a @ b
    torch.cuda.synchronize(dev)
    dt = time.perf_counter() - t0

    flops = 2 * (n ** 3) * iters
    tflops = flops / dt / 1e12
    del a, b, c
    torch.cuda.empty_cache()
    return {"matrix_size": n, "iters": iters, "seconds": round(dt, 4),
            "tflops_bf16": round(tflops, 2)}


def main():
    report = {
        "date_utc": datetime.now(timezone.utc).isoformat(timespec="seconds"),
        "python": platform.python_version(),
        "torch": torch.__version__,
        "torch_cuda_version": torch.version.cuda,
        "cuda_available": torch.cuda.is_available(),
        "device_count": torch.cuda.device_count(),
        "cudnn": torch.backends.cudnn.version(),
        "cuda_device_order": os.environ.get("CUDA_DEVICE_ORDER"),
        "gpus": [],
        "nvidia_smi": nvidia_smi_table(),
    }

    if not report["cuda_available"]:
        print(json.dumps(report, indent=2))
        sys.exit("ECHEC : CUDA indisponible.")

    torch.backends.cuda.matmul.allow_tf32 = True
    torch.backends.cudnn.allow_tf32 = True

    for i in range(torch.cuda.device_count()):
        props = torch.cuda.get_device_properties(i)
        free, total = torch.cuda.mem_get_info(i)
        entry = {
            "torch_index": i,
            "name": props.name,
            "capability": f"{props.major}.{props.minor}",
            "sm_count": props.multi_processor_count,
            "vram_total_mib": total // 2**20,
            "vram_free_mib": free // 2**20,
        }
        print(f"[bench] cuda:{i} {props.name} "
              f"({entry['vram_free_mib']} MiB libres) ...", flush=True)
        entry["matmul_bf16"] = bench_matmul(i)
        print(f"        -> {entry['matmul_bf16']['tflops_bf16']} TFLOPS bf16 "
              f"(matrices {entry['matmul_bf16']['matrix_size']}^2)", flush=True)
        report["gpus"].append(entry)

    out = os.path.join(os.path.dirname(os.path.abspath(__file__)),
                       "logs", "phase0_bench.json")
    with open(out, "w") as f:
        json.dump(report, f, indent=2)
    print(f"\nRapport ecrit dans {out}")
    print(json.dumps(report, indent=2))


if __name__ == "__main__":
    main()
vocab.json3954 lines

The complete vocabulary: the 1971 tokens, ready to copy. It is the project's only data file, and it depends on no corpus.

{
  "itos": [
    "<pad>",
    "<bos>",
    "<eos>",
    "a1a2",
    "a1a3",
    "a1a4",
    "a1a5",
    "a1a6",
    "a1a7",
    "a1a8",
    "a1b1",
    "a1b2",
    "a1b3",
    "a1c1",
    "a1c2",
    "a1c3",
    "a1d1",
    "a1d4",
    "a1e1",
    "a1e5",
    "a1f1",
    "a1f6",
    "a1g1",
    "a1g7",
    "a1h1",
    "a1h8",
    "a2a1",
    "a2a1b",
    "a2a1n",
    "a2a1q",
    "a2a1r",
    "a2a3",
    "a2a4",
    "a2a5",
    "a2a6",
    "a2a7",
    "a2a8",
    "a2b1",
    "a2b1b",
    "a2b1n",
    "a2b1q",
    "a2b1r",
    "a2b2",
    "a2b3",
    "a2b4",
    "a2c1",
    "a2c2",
    "a2c3",
    "a2c4",
    "a2d2",
    "a2d5",
    "a2e2",
    "a2e6",
    "a2f2",
    "a2f7",
    "a2g2",
    "a2g8",
    "a2h2",
    "a3a1",
    "a3a2",
    "a3a4",
    "a3a5",
    "a3a6",
    "a3a7",
    "a3a8",
    "a3b1",
    "a3b2",
    "a3b3",
    "a3b4",
    "a3b5",
    "a3c1",
    "a3c2",
    "a3c3",
    "a3c4",
    "a3c5",
    "a3d3",
    "a3d6",
    "a3e3",
    "a3e7",
    "a3f3",
    "a3f8",
    "a3g3",
    "a3h3",
    "a4a1",
    "a4a2",
    "a4a3",
    "a4a5",
    "a4a6",
    "a4a7",
    "a4a8",
    "a4b2",
    "a4b3",
    "a4b4",
    "a4b5",
    "a4b6",
    "a4c2",
    "a4c3",
    "a4c4",
    "a4c5",
    "a4c6",
    "a4d1",
    "a4d4",
    "a4d7",
    "a4e4",
    "a4e8",
    "a4f4",
    "a4g4",
    "a4h4",
    "a5a1",
    "a5a2",
    "a5a3",
    "a5a4",
    "a5a6",
    "a5a7",
    "a5a8",
    "a5b3",
    "a5b4",
    "a5b5",
    "a5b6",
    "a5b7",
    "a5c3",
    "a5c4",
    "a5c5",
    "a5c6",
    "a5c7",
    "a5d2",
    "a5d5",
    "a5d8",
    "a5e1",
    "a5e5",
    "a5f5",
    "a5g5",
    "a5h5",
    "a6a1",
    "a6a2",
    "a6a3",
    "a6a4",
    "a6a5",
    "a6a7",
    "a6a8",
    "a6b4",
    "a6b5",
    "a6b6",
    "a6b7",
    "a6b8",
    "a6c4",
    "a6c5",
    "a6c6",
    "a6c7",
    "a6c8",
    "a6d3",
    "a6d6",
    "a6e2",
    "a6e6",
    "a6f1",
    "a6f6",
    "a6g6",
    "a6h6",
    "a7a1",
    "a7a2",
    "a7a3",
    "a7a4",
    "a7a5",
    "a7a6",
    "a7a8",
    "a7a8b",
    "a7a8n",
    "a7a8q",
    "a7a8r",
    "a7b5",
    "a7b6",
    "a7b7",
    "a7b8",
    "a7b8b",
    "a7b8n",
    "a7b8q",
    "a7b8r",
    "a7c5",
    "a7c6",
    "a7c7",
    "a7c8",
    "a7d4",
    "a7d7",
    "a7e3",
    "a7e7",
    "a7f2",
    "a7f7",
    "a7g1",
    "a7g7",
    "a7h7",
    "a8a1",
    "a8a2",
    "a8a3",
    "a8a4",
    "a8a5",
    "a8a6",
    "a8a7",
    "a8b6",
    "a8b7",
    "a8b8",
    "a8c6",
    "a8c7",
    "a8c8",
    "a8d5",
    "a8d8",
    "a8e4",
    "a8e8",
    "a8f3",
    "a8f8",
    "a8g2",
    "a8g8",
    "a8h1",
    "a8h8",
    "b1a1",
    "b1a2",
    "b1a3",
    "b1b2",
    "b1b3",
    "b1b4",
    "b1b5",
    "b1b6",
    "b1b7",
    "b1b8",
    "b1c1",
    "b1c2",
    "b1c3",
    "b1d1",
    "b1d2",
    "b1d3",
    "b1e1",
    "b1e4",
    "b1f1",
    "b1f5",
    "b1g1",
    "b1g6",
    "b1h1",
    "b1h7",
    "b2a1",
    "b2a1b",
    "b2a1n",
    "b2a1q",
    "b2a1r",
    "b2a2",
    "b2a3",
    "b2a4",
    "b2b1",
    "b2b1b",
    "b2b1n",
    "b2b1q",
    "b2b1r",
    "b2b3",
    "b2b4",
    "b2b5",
    "b2b6",
    "b2b7",
    "b2b8",
    "b2c1",
    "b2c1b",
    "b2c1n",
    "b2c1q",
    "b2c1r",
    "b2c2",
    "b2c3",
    "b2c4",
    "b2d1",
    "b2d2",
    "b2d3",
    "b2d4",
    "b2e2",
    "b2e5",
    "b2f2",
    "b2f6",
    "b2g2",
    "b2g7",
    "b2h2",
    "b2h8",
    "b3a1",
    "b3a2",
    "b3a3",
    "b3a4",
    "b3a5",
    "b3b1",
    "b3b2",
    "b3b4",
    "b3b5",
    "b3b6",
    "b3b7",
    "b3b8",
    "b3c1",
    "b3c2",
    "b3c3",
    "b3c4",
    "b3c5",
    "b3d1",
    "b3d2",
    "b3d3",
    "b3d4",
    "b3d5",
    "b3e3",
    "b3e6",
    "b3f3",
    "b3f7",
    "b3g3",
    "b3g8",
    "b3h3",
    "b4a2",
    "b4a3",
    "b4a4",
    "b4a5",
    "b4a6",
    "b4b1",
    "b4b2",
    "b4b3",
    "b4b5",
    "b4b6",
    "b4b7",
    "b4b8",
    "b4c2",
    "b4c3",
    "b4c4",
    "b4c5",
    "b4c6",
    "b4d2",
    "b4d3",
    "b4d4",
    "b4d5",
    "b4d6",
    "b4e1",
    "b4e4",
    "b4e7",
    "b4f4",
    "b4f8",
    "b4g4",
    "b4h4",
    "b5a3",
    "b5a4",
    "b5a5",
    "b5a6",
    "b5a7",
    "b5b1",
    "b5b2",
    "b5b3",
    "b5b4",
    "b5b6",
    "b5b7",
    "b5b8",
    "b5c3",
    "b5c4",
    "b5c5",
    "b5c6",
    "b5c7",
    "b5d3",
    "b5d4",
    "b5d5",
    "b5d6",
    "b5d7",
    "b5e2",
    "b5e5",
    "b5e8",
    "b5f1",
    "b5f5",
    "b5g5",
    "b5h5",
    "b6a4",
    "b6a5",
    "b6a6",
    "b6a7",
    "b6a8",
    "b6b1",
    "b6b2",
    "b6b3",
    "b6b4",
    "b6b5",
    "b6b7",
    "b6b8",
    "b6c4",
    "b6c5",
    "b6c6",
    "b6c7",
    "b6c8",
    "b6d4",
    "b6d5",
    "b6d6",
    "b6d7",
    "b6d8",
    "b6e3",
    "b6e6",
    "b6f2",
    "b6f6",
    "b6g1",
    "b6g6",
    "b6h6",
    "b7a5",
    "b7a6",
    "b7a7",
    "b7a8",
    "b7a8b",
    "b7a8n",
    "b7a8q",
    "b7a8r",
    "b7b1",
    "b7b2",
    "b7b3",
    "b7b4",
    "b7b5",
    "b7b6",
    "b7b8",
    "b7b8b",
    "b7b8n",
    "b7b8q",
    "b7b8r",
    "b7c5",
    "b7c6",
    "b7c7",
    "b7c8",
    "b7c8b",
    "b7c8n",
    "b7c8q",
    "b7c8r",
    "b7d5",
    "b7d6",
    "b7d7",
    "b7d8",
    "b7e4",
    "b7e7",
    "b7f3",
    "b7f7",
    "b7g2",
    "b7g7",
    "b7h1",
    "b7h7",
    "b8a6",
    "b8a7",
    "b8a8",
    "b8b1",
    "b8b2",
    "b8b3",
    "b8b4",
    "b8b5",
    "b8b6",
    "b8b7",
    "b8c6",
    "b8c7",
    "b8c8",
    "b8d6",
    "b8d7",
    "b8d8",
    "b8e5",
    "b8e8",
    "b8f4",
    "b8f8",
    "b8g3",
    "b8g8",
    "b8h2",
    "b8h8",
    "c1a1",
    "c1a2",
    "c1a3",
    "c1b1",
    "c1b2",
    "c1b3",
    "c1c2",
    "c1c3",
    "c1c4",
    "c1c5",
    "c1c6",
    "c1c7",
    "c1c8",
    "c1d1",
    "c1d2",
    "c1d3",
    "c1e1",
    "c1e2",
    "c1e3",
    "c1f1",
    "c1f4",
    "c1g1",
    "c1g5",
    "c1h1",
    "c1h6",
    "c2a1",
    "c2a2",
    "c2a3",
    "c2a4",
    "c2b1",
    "c2b1b",
    "c2b1n",
    "c2b1q",
    "c2b1r",
    "c2b2",
    "c2b3",
    "c2b4",
    "c2c1",
    "c2c1b",
    "c2c1n",
    "c2c1q",
    "c2c1r",
    "c2c3",
    "c2c4",
    "c2c5",
    "c2c6",
    "c2c7",
    "c2c8",
    "c2d1",
    "c2d1b",
    "c2d1n",
    "c2d1q",
    "c2d1r",
    "c2d2",
    "c2d3",
    "c2d4",
    "c2e1",
    "c2e2",
    "c2e3",
    "c2e4",
    "c2f2",
    "c2f5",
    "c2g2",
    "c2g6",
    "c2h2",
    "c2h7",
    "c3a1",
    "c3a2",
    "c3a3",
    "c3a4",
    "c3a5",
    "c3b1",
    "c3b2",
    "c3b3",
    "c3b4",
    "c3b5",
    "c3c1",
    "c3c2",
    "c3c4",
    "c3c5",
    "c3c6",
    "c3c7",
    "c3c8",
    "c3d1",
    "c3d2",
    "c3d3",
    "c3d4",
    "c3d5",
    "c3e1",
    "c3e2",
    "c3e3",
    "c3e4",
    "c3e5",
    "c3f3",
    "c3f6",
    "c3g3",
    "c3g7",
    "c3h3",
    "c3h8",
    "c4a2",
    "c4a3",
    "c4a4",
    "c4a5",
    "c4a6",
    "c4b2",
    "c4b3",
    "c4b4",
    "c4b5",
    "c4b6",
    "c4c1",
    "c4c2",
    "c4c3",
    "c4c5",
    "c4c6",
    "c4c7",
    "c4c8",
    "c4d2",
    "c4d3",
    "c4d4",
    "c4d5",
    "c4d6",
    "c4e2",
    "c4e3",
    "c4e4",
    "c4e5",
    "c4e6",
    "c4f1",
    "c4f4",
    "c4f7",
    "c4g4",
    "c4g8",
    "c4h4",
    "c5a3",
    "c5a4",
    "c5a5",
    "c5a6",
    "c5a7",
    "c5b3",
    "c5b4",
    "c5b5",
    "c5b6",
    "c5b7",
    "c5c1",
    "c5c2",
    "c5c3",
    "c5c4",
    "c5c6",
    "c5c7",
    "c5c8",
    "c5d3",
    "c5d4",
    "c5d5",
    "c5d6",
    "c5d7",
    "c5e3",
    "c5e4",
    "c5e5",
    "c5e6",
    "c5e7",
    "c5f2",
    "c5f5",
    "c5f8",
    "c5g1",
    "c5g5",
    "c5h5",
    "c6a4",
    "c6a5",
    "c6a6",
    "c6a7",
    "c6a8",
    "c6b4",
    "c6b5",
    "c6b6",
    "c6b7",
    "c6b8",
    "c6c1",
    "c6c2",
    "c6c3",
    "c6c4",
    "c6c5",
    "c6c7",
    "c6c8",
    "c6d4",
    "c6d5",
    "c6d6",
    "c6d7",
    "c6d8",
    "c6e4",
    "c6e5",
    "c6e6",
    "c6e7",
    "c6e8",
    "c6f3",
    "c6f6",
    "c6g2",
    "c6g6",
    "c6h1",
    "c6h6",
    "c7a5",
    "c7a6",
    "c7a7",
    "c7a8",
    "c7b5",
    "c7b6",
    "c7b7",
    "c7b8",
    "c7b8b",
    "c7b8n",
    "c7b8q",
    "c7b8r",
    "c7c1",
    "c7c2",
    "c7c3",
    "c7c4",
    "c7c5",
    "c7c6",
    "c7c8",
    "c7c8b",
    "c7c8n",
    "c7c8q",
    "c7c8r",
    "c7d5",
    "c7d6",
    "c7d7",
    "c7d8",
    "c7d8b",
    "c7d8n",
    "c7d8q",
    "c7d8r",
    "c7e5",
    "c7e6",
    "c7e7",
    "c7e8",
    "c7f4",
    "c7f7",
    "c7g3",
    "c7g7",
    "c7h2",
    "c7h7",
    "c8a6",
    "c8a7",
    "c8a8",
    "c8b6",
    "c8b7",
    "c8b8",
    "c8c1",
    "c8c2",
    "c8c3",
    "c8c4",
    "c8c5",
    "c8c6",
    "c8c7",
    "c8d6",
    "c8d7",
    "c8d8",
    "c8e6",
    "c8e7",
    "c8e8",
    "c8f5",
    "c8f8",
    "c8g4",
    "c8g8",
    "c8h3",
    "c8h8",
    "d1a1",
    "d1a4",
    "d1b1",
    "d1b2",
    "d1b3",
    "d1c1",
    "d1c2",
    "d1c3",
    "d1d2",
    "d1d3",
    "d1d4",
    "d1d5",
    "d1d6",
    "d1d7",
    "d1d8",
    "d1e1",
    "d1e2",
    "d1e3",
    "d1f1",
    "d1f2",
    "d1f3",
    "d1g1",
    "d1g4",
    "d1h1",
    "d1h5",
    "d2a2",
    "d2a5",
    "d2b1",
    "d2b2",
    "d2b3",
    "d2b4",
    "d2c1",
    "d2c1b",
    "d2c1n",
    "d2c1q",
    "d2c1r",
    "d2c2",
    "d2c3",
    "d2c4",
    "d2d1",
    "d2d1b",
    "d2d1n",
    "d2d1q",
    "d2d1r",
    "d2d3",
    "d2d4",
    "d2d5",
    "d2d6",
    "d2d7",
    "d2d8",
    "d2e1",
    "d2e1b",
    "d2e1n",
    "d2e1q",
    "d2e1r",
    "d2e2",
    "d2e3",
    "d2e4",
    "d2f1",
    "d2f2",
    "d2f3",
    "d2f4",
    "d2g2",
    "d2g5",
    "d2h2",
    "d2h6",
    "d3a3",
    "d3a6",
    "d3b1",
    "d3b2",
    "d3b3",
    "d3b4",
    "d3b5",
    "d3c1",
    "d3c2",
    "d3c3",
    "d3c4",
    "d3c5",
    "d3d1",
    "d3d2",
    "d3d4",
    "d3d5",
    "d3d6",
    "d3d7",
    "d3d8",
    "d3e1",
    "d3e2",
    "d3e3",
    "d3e4",
    "d3e5",
    "d3f1",
    "d3f2",
    "d3f3",
    "d3f4",
    "d3f5",
    "d3g3",
    "d3g6",
    "d3h3",
    "d3h7",
    "d4a1",
    "d4a4",
    "d4a7",
    "d4b2",
    "d4b3",
    "d4b4",
    "d4b5",
    "d4b6",
    "d4c2",
    "d4c3",
    "d4c4",
    "d4c5",
    "d4c6",
    "d4d1",
    "d4d2",
    "d4d3",
    "d4d5",
    "d4d6",
    "d4d7",
    "d4d8",
    "d4e2",
    "d4e3",
    "d4e4",
    "d4e5",
    "d4e6",
    "d4f2",
    "d4f3",
    "d4f4",
    "d4f5",
    "d4f6",
    "d4g1",
    "d4g4",
    "d4g7",
    "d4h4",
    "d4h8",
    "d5a2",
    "d5a5",
    "d5a8",
    "d5b3",
    "d5b4",
    "d5b5",
    "d5b6",
    "d5b7",
    "d5c3",
    "d5c4",
    "d5c5",
    "d5c6",
    "d5c7",
    "d5d1",
    "d5d2",
    "d5d3",
    "d5d4",
    "d5d6",
    "d5d7",
    "d5d8",
    "d5e3",
    "d5e4",
    "d5e5",
    "d5e6",
    "d5e7",
    "d5f3",
    "d5f4",
    "d5f5",
    "d5f6",
    "d5f7",
    "d5g2",
    "d5g5",
    "d5g8",
    "d5h1",
    "d5h5",
    "d6a3",
    "d6a6",
    "d6b4",
    "d6b5",
    "d6b6",
    "d6b7",
    "d6b8",
    "d6c4",
    "d6c5",
    "d6c6",
    "d6c7",
    "d6c8",
    "d6d1",
    "d6d2",
    "d6d3",
    "d6d4",
    "d6d5",
    "d6d7",
    "d6d8",
    "d6e4",
    "d6e5",
    "d6e6",
    "d6e7",
    "d6e8",
    "d6f4",
    "d6f5",
    "d6f6",
    "d6f7",
    "d6f8",
    "d6g3",
    "d6g6",
    "d6h2",
    "d6h6",
    "d7a4",
    "d7a7",
    "d7b5",
    "d7b6",
    "d7b7",
    "d7b8",
    "d7c5",
    "d7c6",
    "d7c7",
    "d7c8",
    "d7c8b",
    "d7c8n",
    "d7c8q",
    "d7c8r",
    "d7d1",
    "d7d2",
    "d7d3",
    "d7d4",
    "d7d5",
    "d7d6",
    "d7d8",
    "d7d8b",
    "d7d8n",
    "d7d8q",
    "d7d8r",
    "d7e5",
    "d7e6",
    "d7e7",
    "d7e8",
    "d7e8b",
    "d7e8n",
    "d7e8q",
    "d7e8r",
    "d7f5",
    "d7f6",
    "d7f7",
    "d7f8",
    "d7g4",
    "d7g7",
    "d7h3",
    "d7h7",
    "d8a5",
    "d8a8",
    "d8b6",
    "d8b7",
    "d8b8",
    "d8c6",
    "d8c7",
    "d8c8",
    "d8d1",
    "d8d2",
    "d8d3",
    "d8d4",
    "d8d5",
    "d8d6",
    "d8d7",
    "d8e6",
    "d8e7",
    "d8e8",
    "d8f6",
    "d8f7",
    "d8f8",
    "d8g5",
    "d8g8",
    "d8h4",
    "d8h8",
    "e1a1",
    "e1a5",
    "e1b1",
    "e1b4",
    "e1c1",
    "e1c2",
    "e1c3",
    "e1d1",
    "e1d2",
    "e1d3",
    "e1e2",
    "e1e3",
    "e1e4",
    "e1e5",
    "e1e6",
    "e1e7",
    "e1e8",
    "e1f1",
    "e1f2",
    "e1f3",
    "e1g1",
    "e1g2",
    "e1g3",
    "e1h1",
    "e1h4",
    "e2a2",
    "e2a6",
    "e2b2",
    "e2b5",
    "e2c1",
    "e2c2",
    "e2c3",
    "e2c4",
    "e2d1",
    "e2d1b",
    "e2d1n",
    "e2d1q",
    "e2d1r",
    "e2d2",
    "e2d3",
    "e2d4",
    "e2e1",
    "e2e1b",
    "e2e1n",
    "e2e1q",
    "e2e1r",
    "e2e3",
    "e2e4",
    "e2e5",
    "e2e6",
    "e2e7",
    "e2e8",
    "e2f1",
    "e2f1b",
    "e2f1n",
    "e2f1q",
    "e2f1r",
    "e2f2",
    "e2f3",
    "e2f4",
    "e2g1",
    "e2g2",
    "e2g3",
    "e2g4",
    "e2h2",
    "e2h5",
    "e3a3",
    "e3a7",
    "e3b3",
    "e3b6",
    "e3c1",
    "e3c2",
    "e3c3",
    "e3c4",
    "e3c5",
    "e3d1",
    "e3d2",
    "e3d3",
    "e3d4",
    "e3d5",
    "e3e1",
    "e3e2",
    "e3e4",
    "e3e5",
    "e3e6",
    "e3e7",
    "e3e8",
    "e3f1",
    "e3f2",
    "e3f3",
    "e3f4",
    "e3f5",
    "e3g1",
    "e3g2",
    "e3g3",
    "e3g4",
    "e3g5",
    "e3h3",
    "e3h6",
    "e4a4",
    "e4a8",
    "e4b1",
    "e4b4",
    "e4b7",
    "e4c2",
    "e4c3",
    "e4c4",
    "e4c5",
    "e4c6",
    "e4d2",
    "e4d3",
    "e4d4",
    "e4d5",
    "e4d6",
    "e4e1",
    "e4e2",
    "e4e3",
    "e4e5",
    "e4e6",
    "e4e7",
    "e4e8",
    "e4f2",
    "e4f3",
    "e4f4",
    "e4f5",
    "e4f6",
    "e4g2",
    "e4g3",
    "e4g4",
    "e4g5",
    "e4g6",
    "e4h1",
    "e4h4",
    "e4h7",
    "e5a1",
    "e5a5",
    "e5b2",
    "e5b5",
    "e5b8",
    "e5c3",
    "e5c4",
    "e5c5",
    "e5c6",
    "e5c7",
    "e5d3",
    "e5d4",
    "e5d5",
    "e5d6",
    "e5d7",
    "e5e1",
    "e5e2",
    "e5e3",
    "e5e4",
    "e5e6",
    "e5e7",
    "e5e8",
    "e5f3",
    "e5f4",
    "e5f5",
    "e5f6",
    "e5f7",
    "e5g3",
    "e5g4",
    "e5g5",
    "e5g6",
    "e5g7",
    "e5h2",
    "e5h5",
    "e5h8",
    "e6a2",
    "e6a6",
    "e6b3",
    "e6b6",
    "e6c4",
    "e6c5",
    "e6c6",
    "e6c7",
    "e6c8",
    "e6d4",
    "e6d5",
    "e6d6",
    "e6d7",
    "e6d8",
    "e6e1",
    "e6e2",
    "e6e3",
    "e6e4",
    "e6e5",
    "e6e7",
    "e6e8",
    "e6f4",
    "e6f5",
    "e6f6",
    "e6f7",
    "e6f8",
    "e6g4",
    "e6g5",
    "e6g6",
    "e6g7",
    "e6g8",
    "e6h3",
    "e6h6",
    "e7a3",
    "e7a7",
    "e7b4",
    "e7b7",
    "e7c5",
    "e7c6",
    "e7c7",
    "e7c8",
    "e7d5",
    "e7d6",
    "e7d7",
    "e7d8",
    "e7d8b",
    "e7d8n",
    "e7d8q",
    "e7d8r",
    "e7e1",
    "e7e2",
    "e7e3",
    "e7e4",
    "e7e5",
    "e7e6",
    "e7e8",
    "e7e8b",
    "e7e8n",
    "e7e8q",
    "e7e8r",
    "e7f5",
    "e7f6",
    "e7f7",
    "e7f8",
    "e7f8b",
    "e7f8n",
    "e7f8q",
    "e7f8r",
    "e7g5",
    "e7g6",
    "e7g7",
    "e7g8",
    "e7h4",
    "e7h7",
    "e8a4",
    "e8a8",
    "e8b5",
    "e8b8",
    "e8c6",
    "e8c7",
    "e8c8",
    "e8d6",
    "e8d7",
    "e8d8",
    "e8e1",
    "e8e2",
    "e8e3",
    "e8e4",
    "e8e5",
    "e8e6",
    "e8e7",
    "e8f6",
    "e8f7",
    "e8f8",
    "e8g6",
    "e8g7",
    "e8g8",
    "e8h5",
    "e8h8",
    "f1a1",
    "f1a6",
    "f1b1",
    "f1b5",
    "f1c1",
    "f1c4",
    "f1d1",
    "f1d2",
    "f1d3",
    "f1e1",
    "f1e2",
    "f1e3",
    "f1f2",
    "f1f3",
    "f1f4",
    "f1f5",
    "f1f6",
    "f1f7",
    "f1f8",
    "f1g1",
    "f1g2",
    "f1g3",
    "f1h1",
    "f1h2",
    "f1h3",
    "f2a2",
    "f2a7",
    "f2b2",
    "f2b6",
    "f2c2",
    "f2c5",
    "f2d1",
    "f2d2",
    "f2d3",
    "f2d4",
    "f2e1",
    "f2e1b",
    "f2e1n",
    "f2e1q",
    "f2e1r",
    "f2e2",
    "f2e3",
    "f2e4",
    "f2f1",
    "f2f1b",
    "f2f1n",
    "f2f1q",
    "f2f1r",
    "f2f3",
    "f2f4",
    "f2f5",
    "f2f6",
    "f2f7",
    "f2f8",
    "f2g1",
    "f2g1b",
    "f2g1n",
    "f2g1q",
    "f2g1r",
    "f2g2",
    "f2g3",
    "f2g4",
    "f2h1",
    "f2h2",
    "f2h3",
    "f2h4",
    "f3a3",
    "f3a8",
    "f3b3",
    "f3b7",
    "f3c3",
    "f3c6",
    "f3d1",
    "f3d2",
    "f3d3",
    "f3d4",
    "f3d5",
    "f3e1",
    "f3e2",
    "f3e3",
    "f3e4",
    "f3e5",
    "f3f1",
    "f3f2",
    "f3f4",
    "f3f5",
    "f3f6",
    "f3f7",
    "f3f8",
    "f3g1",
    "f3g2",
    "f3g3",
    "f3g4",
    "f3g5",
    "f3h1",
    "f3h2",
    "f3h3",
    "f3h4",
    "f3h5",
    "f4a4",
    "f4b4",
    "f4b8",
    "f4c1",
    "f4c4",
    "f4c7",
    "f4d2",
    "f4d3",
    "f4d4",
    "f4d5",
    "f4d6",
    "f4e2",
    "f4e3",
    "f4e4",
    "f4e5",
    "f4e6",
    "f4f1",
    "f4f2",
    "f4f3",
    "f4f5",
    "f4f6",
    "f4f7",
    "f4f8",
    "f4g2",
    "f4g3",
    "f4g4",
    "f4g5",
    "f4g6",
    "f4h2",
    "f4h3",
    "f4h4",
    "f4h5",
    "f4h6",
    "f5a5",
    "f5b1",
    "f5b5",
    "f5c2",
    "f5c5",
    "f5c8",
    "f5d3",
    "f5d4",
    "f5d5",
    "f5d6",
    "f5d7",
    "f5e3",
    "f5e4",
    "f5e5",
    "f5e6",
    "f5e7",
    "f5f1",
    "f5f2",
    "f5f3",
    "f5f4",
    "f5f6",
    "f5f7",
    "f5f8",
    "f5g3",
    "f5g4",
    "f5g5",
    "f5g6",
    "f5g7",
    "f5h3",
    "f5h4",
    "f5h5",
    "f5h6",
    "f5h7",
    "f6a1",
    "f6a6",
    "f6b2",
    "f6b6",
    "f6c3",
    "f6c6",
    "f6d4",
    "f6d5",
    "f6d6",
    "f6d7",
    "f6d8",
    "f6e4",
    "f6e5",
    "f6e6",
    "f6e7",
    "f6e8",
    "f6f1",
    "f6f2",
    "f6f3",
    "f6f4",
    "f6f5",
    "f6f7",
    "f6f8",
    "f6g4",
    "f6g5",
    "f6g6",
    "f6g7",
    "f6g8",
    "f6h4",
    "f6h5",
    "f6h6",
    "f6h7",
    "f6h8",
    "f7a2",
    "f7a7",
    "f7b3",
    "f7b7",
    "f7c4",
    "f7c7",
    "f7d5",
    "f7d6",
    "f7d7",
    "f7d8",
    "f7e5",
    "f7e6",
    "f7e7",
    "f7e8",
    "f7e8b",
    "f7e8n",
    "f7e8q",
    "f7e8r",
    "f7f1",
    "f7f2",
    "f7f3",
    "f7f4",
    "f7f5",
    "f7f6",
    "f7f8",
    "f7f8b",
    "f7f8n",
    "f7f8q",
    "f7f8r",
    "f7g5",
    "f7g6",
    "f7g7",
    "f7g8",
    "f7g8b",
    "f7g8n",
    "f7g8q",
    "f7g8r",
    "f7h5",
    "f7h6",
    "f7h7",
    "f7h8",
    "f8a3",
    "f8a8",
    "f8b4",
    "f8b8",
    "f8c5",
    "f8c8",
    "f8d6",
    "f8d7",
    "f8d8",
    "f8e6",
    "f8e7",
    "f8e8",
    "f8f1",
    "f8f2",
    "f8f3",
    "f8f4",
    "f8f5",
    "f8f6",
    "f8f7",
    "f8g6",
    "f8g7",
    "f8g8",
    "f8h6",
    "f8h7",
    "f8h8",
    "g1a1",
    "g1a7",
    "g1b1",
    "g1b6",
    "g1c1",
    "g1c5",
    "g1d1",
    "g1d4",
    "g1e1",
    "g1e2",
    "g1e3",
    "g1f1",
    "g1f2",
    "g1f3",
    "g1g2",
    "g1g3",
    "g1g4",
    "g1g5",
    "g1g6",
    "g1g7",
    "g1g8",
    "g1h1",
    "g1h2",
    "g1h3",
    "g2a2",
    "g2a8",
    "g2b2",
    "g2b7",
    "g2c2",
    "g2c6",
    "g2d2",
    "g2d5",
    "g2e1",
    "g2e2",
    "g2e3",
    "g2e4",
    "g2f1",
    "g2f1b",
    "g2f1n",
    "g2f1q",
    "g2f1r",
    "g2f2",
    "g2f3",
    "g2f4",
    "g2g1",
    "g2g1b",
    "g2g1n",
    "g2g1q",
    "g2g1r",
    "g2g3",
    "g2g4",
    "g2g5",
    "g2g6",
    "g2g7",
    "g2g8",
    "g2h1",
    "g2h1b",
    "g2h1n",
    "g2h1q",
    "g2h1r",
    "g2h2",
    "g2h3",
    "g2h4",
    "g3a3",
    "g3b3",
    "g3b8",
    "g3c3",
    "g3c7",
    "g3d3",
    "g3d6",
    "g3e1",
    "g3e2",
    "g3e3",
    "g3e4",
    "g3e5",
    "g3f1",
    "g3f2",
    "g3f3",
    "g3f4",
    "g3f5",
    "g3g1",
    "g3g2",
    "g3g4",
    "g3g5",
    "g3g6",
    "g3g7",
    "g3g8",
    "g3h1",
    "g3h2",
    "g3h3",
    "g3h4",
    "g3h5",
    "g4a4",
    "g4b4",
    "g4c4",
    "g4c8",
    "g4d1",
    "g4d4",
    "g4d7",
    "g4e2",
    "g4e3",
    "g4e4",
    "g4e5",
    "g4e6",
    "g4f2",
    "g4f3",
    "g4f4",
    "g4f5",
    "g4f6",
    "g4g1",
    "g4g2",
    "g4g3",
    "g4g5",
    "g4g6",
    "g4g7",
    "g4g8",
    "g4h2",
    "g4h3",
    "g4h4",
    "g4h5",
    "g4h6",
    "g5a5",
    "g5b5",
    "g5c1",
    "g5c5",
    "g5d2",
    "g5d5",
    "g5d8",
    "g5e3",
    "g5e4",
    "g5e5",
    "g5e6",
    "g5e7",
    "g5f3",
    "g5f4",
    "g5f5",
    "g5f6",
    "g5f7",
    "g5g1",
    "g5g2",
    "g5g3",
    "g5g4",
    "g5g6",
    "g5g7",
    "g5g8",
    "g5h3",
    "g5h4",
    "g5h5",
    "g5h6",
    "g5h7",
    "g6a6",
    "g6b1",
    "g6b6",
    "g6c2",
    "g6c6",
    "g6d3",
    "g6d6",
    "g6e4",
    "g6e5",
    "g6e6",
    "g6e7",
    "g6e8",
    "g6f4",
    "g6f5",
    "g6f6",
    "g6f7",
    "g6f8",
    "g6g1",
    "g6g2",
    "g6g3",
    "g6g4",
    "g6g5",
    "g6g7",
    "g6g8",
    "g6h4",
    "g6h5",
    "g6h6",
    "g6h7",
    "g6h8",
    "g7a1",
    "g7a7",
    "g7b2",
    "g7b7",
    "g7c3",
    "g7c7",
    "g7d4",
    "g7d7",
    "g7e5",
    "g7e6",
    "g7e7",
    "g7e8",
    "g7f5",
    "g7f6",
    "g7f7",
    "g7f8",
    "g7f8b",
    "g7f8n",
    "g7f8q",
    "g7f8r",
    "g7g1",
    "g7g2",
    "g7g3",
    "g7g4",
    "g7g5",
    "g7g6",
    "g7g8",
    "g7g8b",
    "g7g8n",
    "g7g8q",
    "g7g8r",
    "g7h5",
    "g7h6",
    "g7h7",
    "g7h8",
    "g7h8b",
    "g7h8n",
    "g7h8q",
    "g7h8r",
    "g8a2",
    "g8a8",
    "g8b3",
    "g8b8",
    "g8c4",
    "g8c8",
    "g8d5",
    "g8d8",
    "g8e6",
    "g8e7",
    "g8e8",
    "g8f6",
    "g8f7",
    "g8f8",
    "g8g1",
    "g8g2",
    "g8g3",
    "g8g4",
    "g8g5",
    "g8g6",
    "g8g7",
    "g8h6",
    "g8h7",
    "g8h8",
    "h1a1",
    "h1a8",
    "h1b1",
    "h1b7",
    "h1c1",
    "h1c6",
    "h1d1",
    "h1d5",
    "h1e1",
    "h1e4",
    "h1f1",
    "h1f2",
    "h1f3",
    "h1g1",
    "h1g2",
    "h1g3",
    "h1h2",
    "h1h3",
    "h1h4",
    "h1h5",
    "h1h6",
    "h1h7",
    "h1h8",
    "h2a2",
    "h2b2",
    "h2b8",
    "h2c2",
    "h2c7",
    "h2d2",
    "h2d6",
    "h2e2",
    "h2e5",
    "h2f1",
    "h2f2",
    "h2f3",
    "h2f4",
    "h2g1",
    "h2g1b",
    "h2g1n",
    "h2g1q",
    "h2g1r",
    "h2g2",
    "h2g3",
    "h2g4",
    "h2h1",
    "h2h1b",
    "h2h1n",
    "h2h1q",
    "h2h1r",
    "h2h3",
    "h2h4",
    "h2h5",
    "h2h6",
    "h2h7",
    "h2h8",
    "h3a3",
    "h3b3",
    "h3c3",
    "h3c8",
    "h3d3",
    "h3d7",
    "h3e3",
    "h3e6",
    "h3f1",
    "h3f2",
    "h3f3",
    "h3f4",
    "h3f5",
    "h3g1",
    "h3g2",
    "h3g3",
    "h3g4",
    "h3g5",
    "h3h1",
    "h3h2",
    "h3h4",
    "h3h5",
    "h3h6",
    "h3h7",
    "h3h8",
    "h4a4",
    "h4b4",
    "h4c4",
    "h4d4",
    "h4d8",
    "h4e1",
    "h4e4",
    "h4e7",
    "h4f2",
    "h4f3",
    "h4f4",
    "h4f5",
    "h4f6",
    "h4g2",
    "h4g3",
    "h4g4",
    "h4g5",
    "h4g6",
    "h4h1",
    "h4h2",
    "h4h3",
    "h4h5",
    "h4h6",
    "h4h7",
    "h4h8",
    "h5a5",
    "h5b5",
    "h5c5",
    "h5d1",
    "h5d5",
    "h5e2",
    "h5e5",
    "h5e8",
    "h5f3",
    "h5f4",
    "h5f5",
    "h5f6",
    "h5f7",
    "h5g3",
    "h5g4",
    "h5g5",
    "h5g6",
    "h5g7",
    "h5h1",
    "h5h2",
    "h5h3",
    "h5h4",
    "h5h6",
    "h5h7",
    "h5h8",
    "h6a6",
    "h6b6",
    "h6c1",
    "h6c6",
    "h6d2",
    "h6d6",
    "h6e3",
    "h6e6",
    "h6f4",
    "h6f5",
    "h6f6",
    "h6f7",
    "h6f8",
    "h6g4",
    "h6g5",
    "h6g6",
    "h6g7",
    "h6g8",
    "h6h1",
    "h6h2",
    "h6h3",
    "h6h4",
    "h6h5",
    "h6h7",
    "h6h8",
    "h7a7",
    "h7b1",
    "h7b7",
    "h7c2",
    "h7c7",
    "h7d3",
    "h7d7",
    "h7e4",
    "h7e7",
    "h7f5",
    "h7f6",
    "h7f7",
    "h7f8",
    "h7g5",
    "h7g6",
    "h7g7",
    "h7g8",
    "h7g8b",
    "h7g8n",
    "h7g8q",
    "h7g8r",
    "h7h1",
    "h7h2",
    "h7h3",
    "h7h4",
    "h7h5",
    "h7h6",
    "h7h8",
    "h7h8b",
    "h7h8n",
    "h7h8q",
    "h7h8r",
    "h8a1",
    "h8a8",
    "h8b2",
    "h8b8",
    "h8c3",
    "h8c8",
    "h8d4",
    "h8d8",
    "h8e5",
    "h8e8",
    "h8f6",
    "h8f7",
    "h8f8",
    "h8g6",
    "h8g7",
    "h8g8",
    "h8h1",
    "h8h2",
    "h8h3",
    "h8h4",
    "h8h5",
    "h8h6",
    "h8h7"
  ],
  "stoi": {
    "<pad>": 0,
    "<bos>": 1,
    "<eos>": 2,
    "a1a2": 3,
    "a1a3": 4,
    "a1a4": 5,
    "a1a5": 6,
    "a1a6": 7,
    "a1a7": 8,
    "a1a8": 9,
    "a1b1": 10,
    "a1b2": 11,
    "a1b3": 12,
    "a1c1": 13,
    "a1c2": 14,
    "a1c3": 15,
    "a1d1": 16,
    "a1d4": 17,
    "a1e1": 18,
    "a1e5": 19,
    "a1f1": 20,
    "a1f6": 21,
    "a1g1": 22,
    "a1g7": 23,
    "a1h1": 24,
    "a1h8": 25,
    "a2a1": 26,
    "a2a1b": 27,
    "a2a1n": 28,
    "a2a1q": 29,
    "a2a1r": 30,
    "a2a3": 31,
    "a2a4": 32,
    "a2a5": 33,
    "a2a6": 34,
    "a2a7": 35,
    "a2a8": 36,
    "a2b1": 37,
    "a2b1b": 38,
    "a2b1n": 39,
    "a2b1q": 40,
    "a2b1r": 41,
    "a2b2": 42,
    "a2b3": 43,
    "a2b4": 44,
    "a2c1": 45,
    "a2c2": 46,
    "a2c3": 47,
    "a2c4": 48,
    "a2d2": 49,
    "a2d5": 50,
    "a2e2": 51,
    "a2e6": 52,
    "a2f2": 53,
    "a2f7": 54,
    "a2g2": 55,
    "a2g8": 56,
    "a2h2": 57,
    "a3a1": 58,
    "a3a2": 59,
    "a3a4": 60,
    "a3a5": 61,
    "a3a6": 62,
    "a3a7": 63,
    "a3a8": 64,
    "a3b1": 65,
    "a3b2": 66,
    "a3b3": 67,
    "a3b4": 68,
    "a3b5": 69,
    "a3c1": 70,
    "a3c2": 71,
    "a3c3": 72,
    "a3c4": 73,
    "a3c5": 74,
    "a3d3": 75,
    "a3d6": 76,
    "a3e3": 77,
    "a3e7": 78,
    "a3f3": 79,
    "a3f8": 80,
    "a3g3": 81,
    "a3h3": 82,
    "a4a1": 83,
    "a4a2": 84,
    "a4a3": 85,
    "a4a5": 86,
    "a4a6": 87,
    "a4a7": 88,
    "a4a8": 89,
    "a4b2": 90,
    "a4b3": 91,
    "a4b4": 92,
    "a4b5": 93,
    "a4b6": 94,
    "a4c2": 95,
    "a4c3": 96,
    "a4c4": 97,
    "a4c5": 98,
    "a4c6": 99,
    "a4d1": 100,
    "a4d4": 101,
    "a4d7": 102,
    "a4e4": 103,
    "a4e8": 104,
    "a4f4": 105,
    "a4g4": 106,
    "a4h4": 107,
    "a5a1": 108,
    "a5a2": 109,
    "a5a3": 110,
    "a5a4": 111,
    "a5a6": 112,
    "a5a7": 113,
    "a5a8": 114,
    "a5b3": 115,
    "a5b4": 116,
    "a5b5": 117,
    "a5b6": 118,
    "a5b7": 119,
    "a5c3": 120,
    "a5c4": 121,
    "a5c5": 122,
    "a5c6": 123,
    "a5c7": 124,
    "a5d2": 125,
    "a5d5": 126,
    "a5d8": 127,
    "a5e1": 128,
    "a5e5": 129,
    "a5f5": 130,
    "a5g5": 131,
    "a5h5": 132,
    "a6a1": 133,
    "a6a2": 134,
    "a6a3": 135,
    "a6a4": 136,
    "a6a5": 137,
    "a6a7": 138,
    "a6a8": 139,
    "a6b4": 140,
    "a6b5": 141,
    "a6b6": 142,
    "a6b7": 143,
    "a6b8": 144,
    "a6c4": 145,
    "a6c5": 146,
    "a6c6": 147,
    "a6c7": 148,
    "a6c8": 149,
    "a6d3": 150,
    "a6d6": 151,
    "a6e2": 152,
    "a6e6": 153,
    "a6f1": 154,
    "a6f6": 155,
    "a6g6": 156,
    "a6h6": 157,
    "a7a1": 158,
    "a7a2": 159,
    "a7a3": 160,
    "a7a4": 161,
    "a7a5": 162,
    "a7a6": 163,
    "a7a8": 164,
    "a7a8b": 165,
    "a7a8n": 166,
    "a7a8q": 167,
    "a7a8r": 168,
    "a7b5": 169,
    "a7b6": 170,
    "a7b7": 171,
    "a7b8": 172,
    "a7b8b": 173,
    "a7b8n": 174,
    "a7b8q": 175,
    "a7b8r": 176,
    "a7c5": 177,
    "a7c6": 178,
    "a7c7": 179,
    "a7c8": 180,
    "a7d4": 181,
    "a7d7": 182,
    "a7e3": 183,
    "a7e7": 184,
    "a7f2": 185,
    "a7f7": 186,
    "a7g1": 187,
    "a7g7": 188,
    "a7h7": 189,
    "a8a1": 190,
    "a8a2": 191,
    "a8a3": 192,
    "a8a4": 193,
    "a8a5": 194,
    "a8a6": 195,
    "a8a7": 196,
    "a8b6": 197,
    "a8b7": 198,
    "a8b8": 199,
    "a8c6": 200,
    "a8c7": 201,
    "a8c8": 202,
    "a8d5": 203,
    "a8d8": 204,
    "a8e4": 205,
    "a8e8": 206,
    "a8f3": 207,
    "a8f8": 208,
    "a8g2": 209,
    "a8g8": 210,
    "a8h1": 211,
    "a8h8": 212,
    "b1a1": 213,
    "b1a2": 214,
    "b1a3": 215,
    "b1b2": 216,
    "b1b3": 217,
    "b1b4": 218,
    "b1b5": 219,
    "b1b6": 220,
    "b1b7": 221,
    "b1b8": 222,
    "b1c1": 223,
    "b1c2": 224,
    "b1c3": 225,
    "b1d1": 226,
    "b1d2": 227,
    "b1d3": 228,
    "b1e1": 229,
    "b1e4": 230,
    "b1f1": 231,
    "b1f5": 232,
    "b1g1": 233,
    "b1g6": 234,
    "b1h1": 235,
    "b1h7": 236,
    "b2a1": 237,
    "b2a1b": 238,
    "b2a1n": 239,
    "b2a1q": 240,
    "b2a1r": 241,
    "b2a2": 242,
    "b2a3": 243,
    "b2a4": 244,
    "b2b1": 245,
    "b2b1b": 246,
    "b2b1n": 247,
    "b2b1q": 248,
    "b2b1r": 249,
    "b2b3": 250,
    "b2b4": 251,
    "b2b5": 252,
    "b2b6": 253,
    "b2b7": 254,
    "b2b8": 255,
    "b2c1": 256,
    "b2c1b": 257,
    "b2c1n": 258,
    "b2c1q": 259,
    "b2c1r": 260,
    "b2c2": 261,
    "b2c3": 262,
    "b2c4": 263,
    "b2d1": 264,
    "b2d2": 265,
    "b2d3": 266,
    "b2d4": 267,
    "b2e2": 268,
    "b2e5": 269,
    "b2f2": 270,
    "b2f6": 271,
    "b2g2": 272,
    "b2g7": 273,
    "b2h2": 274,
    "b2h8": 275,
    "b3a1": 276,
    "b3a2": 277,
    "b3a3": 278,
    "b3a4": 279,
    "b3a5": 280,
    "b3b1": 281,
    "b3b2": 282,
    "b3b4": 283,
    "b3b5": 284,
    "b3b6": 285,
    "b3b7": 286,
    "b3b8": 287,
    "b3c1": 288,
    "b3c2": 289,
    "b3c3": 290,
    "b3c4": 291,
    "b3c5": 292,
    "b3d1": 293,
    "b3d2": 294,
    "b3d3": 295,
    "b3d4": 296,
    "b3d5": 297,
    "b3e3": 298,
    "b3e6": 299,
    "b3f3": 300,
    "b3f7": 301,
    "b3g3": 302,
    "b3g8": 303,
    "b3h3": 304,
    "b4a2": 305,
    "b4a3": 306,
    "b4a4": 307,
    "b4a5": 308,
    "b4a6": 309,
    "b4b1": 310,
    "b4b2": 311,
    "b4b3": 312,
    "b4b5": 313,
    "b4b6": 314,
    "b4b7": 315,
    "b4b8": 316,
    "b4c2": 317,
    "b4c3": 318,
    "b4c4": 319,
    "b4c5": 320,
    "b4c6": 321,
    "b4d2": 322,
    "b4d3": 323,
    "b4d4": 324,
    "b4d5": 325,
    "b4d6": 326,
    "b4e1": 327,
    "b4e4": 328,
    "b4e7": 329,
    "b4f4": 330,
    "b4f8": 331,
    "b4g4": 332,
    "b4h4": 333,
    "b5a3": 334,
    "b5a4": 335,
    "b5a5": 336,
    "b5a6": 337,
    "b5a7": 338,
    "b5b1": 339,
    "b5b2": 340,
    "b5b3": 341,
    "b5b4": 342,
    "b5b6": 343,
    "b5b7": 344,
    "b5b8": 345,
    "b5c3": 346,
    "b5c4": 347,
    "b5c5": 348,
    "b5c6": 349,
    "b5c7": 350,
    "b5d3": 351,
    "b5d4": 352,
    "b5d5": 353,
    "b5d6": 354,
    "b5d7": 355,
    "b5e2": 356,
    "b5e5": 357,
    "b5e8": 358,
    "b5f1": 359,
    "b5f5": 360,
    "b5g5": 361,
    "b5h5": 362,
    "b6a4": 363,
    "b6a5": 364,
    "b6a6": 365,
    "b6a7": 366,
    "b6a8": 367,
    "b6b1": 368,
    "b6b2": 369,
    "b6b3": 370,
    "b6b4": 371,
    "b6b5": 372,
    "b6b7": 373,
    "b6b8": 374,
    "b6c4": 375,
    "b6c5": 376,
    "b6c6": 377,
    "b6c7": 378,
    "b6c8": 379,
    "b6d4": 380,
    "b6d5": 381,
    "b6d6": 382,
    "b6d7": 383,
    "b6d8": 384,
    "b6e3": 385,
    "b6e6": 386,
    "b6f2": 387,
    "b6f6": 388,
    "b6g1": 389,
    "b6g6": 390,
    "b6h6": 391,
    "b7a5": 392,
    "b7a6": 393,
    "b7a7": 394,
    "b7a8": 395,
    "b7a8b": 396,
    "b7a8n": 397,
    "b7a8q": 398,
    "b7a8r": 399,
    "b7b1": 400,
    "b7b2": 401,
    "b7b3": 402,
    "b7b4": 403,
    "b7b5": 404,
    "b7b6": 405,
    "b7b8": 406,
    "b7b8b": 407,
    "b7b8n": 408,
    "b7b8q": 409,
    "b7b8r": 410,
    "b7c5": 411,
    "b7c6": 412,
    "b7c7": 413,
    "b7c8": 414,
    "b7c8b": 415,
    "b7c8n": 416,
    "b7c8q": 417,
    "b7c8r": 418,
    "b7d5": 419,
    "b7d6": 420,
    "b7d7": 421,
    "b7d8": 422,
    "b7e4": 423,
    "b7e7": 424,
    "b7f3": 425,
    "b7f7": 426,
    "b7g2": 427,
    "b7g7": 428,
    "b7h1": 429,
    "b7h7": 430,
    "b8a6": 431,
    "b8a7": 432,
    "b8a8": 433,
    "b8b1": 434,
    "b8b2": 435,
    "b8b3": 436,
    "b8b4": 437,
    "b8b5": 438,
    "b8b6": 439,
    "b8b7": 440,
    "b8c6": 441,
    "b8c7": 442,
    "b8c8": 443,
    "b8d6": 444,
    "b8d7": 445,
    "b8d8": 446,
    "b8e5": 447,
    "b8e8": 448,
    "b8f4": 449,
    "b8f8": 450,
    "b8g3": 451,
    "b8g8": 452,
    "b8h2": 453,
    "b8h8": 454,
    "c1a1": 455,
    "c1a2": 456,
    "c1a3": 457,
    "c1b1": 458,
    "c1b2": 459,
    "c1b3": 460,
    "c1c2": 461,
    "c1c3": 462,
    "c1c4": 463,
    "c1c5": 464,
    "c1c6": 465,
    "c1c7": 466,
    "c1c8": 467,
    "c1d1": 468,
    "c1d2": 469,
    "c1d3": 470,
    "c1e1": 471,
    "c1e2": 472,
    "c1e3": 473,
    "c1f1": 474,
    "c1f4": 475,
    "c1g1": 476,
    "c1g5": 477,
    "c1h1": 478,
    "c1h6": 479,
    "c2a1": 480,
    "c2a2": 481,
    "c2a3": 482,
    "c2a4": 483,
    "c2b1": 484,
    "c2b1b": 485,
    "c2b1n": 486,
    "c2b1q": 487,
    "c2b1r": 488,
    "c2b2": 489,
    "c2b3": 490,
    "c2b4": 491,
    "c2c1": 492,
    "c2c1b": 493,
    "c2c1n": 494,
    "c2c1q": 495,
    "c2c1r": 496,
    "c2c3": 497,
    "c2c4": 498,
    "c2c5": 499,
    "c2c6": 500,
    "c2c7": 501,
    "c2c8": 502,
    "c2d1": 503,
    "c2d1b": 504,
    "c2d1n": 505,
    "c2d1q": 506,
    "c2d1r": 507,
    "c2d2": 508,
    "c2d3": 509,
    "c2d4": 510,
    "c2e1": 511,
    "c2e2": 512,
    "c2e3": 513,
    "c2e4": 514,
    "c2f2": 515,
    "c2f5": 516,
    "c2g2": 517,
    "c2g6": 518,
    "c2h2": 519,
    "c2h7": 520,
    "c3a1": 521,
    "c3a2": 522,
    "c3a3": 523,
    "c3a4": 524,
    "c3a5": 525,
    "c3b1": 526,
    "c3b2": 527,
    "c3b3": 528,
    "c3b4": 529,
    "c3b5": 530,
    "c3c1": 531,
    "c3c2": 532,
    "c3c4": 533,
    "c3c5": 534,
    "c3c6": 535,
    "c3c7": 536,
    "c3c8": 537,
    "c3d1": 538,
    "c3d2": 539,
    "c3d3": 540,
    "c3d4": 541,
    "c3d5": 542,
    "c3e1": 543,
    "c3e2": 544,
    "c3e3": 545,
    "c3e4": 546,
    "c3e5": 547,
    "c3f3": 548,
    "c3f6": 549,
    "c3g3": 550,
    "c3g7": 551,
    "c3h3": 552,
    "c3h8": 553,
    "c4a2": 554,
    "c4a3": 555,
    "c4a4": 556,
    "c4a5": 557,
    "c4a6": 558,
    "c4b2": 559,
    "c4b3": 560,
    "c4b4": 561,
    "c4b5": 562,
    "c4b6": 563,
    "c4c1": 564,
    "c4c2": 565,
    "c4c3": 566,
    "c4c5": 567,
    "c4c6": 568,
    "c4c7": 569,
    "c4c8": 570,
    "c4d2": 571,
    "c4d3": 572,
    "c4d4": 573,
    "c4d5": 574,
    "c4d6": 575,
    "c4e2": 576,
    "c4e3": 577,
    "c4e4": 578,
    "c4e5": 579,
    "c4e6": 580,
    "c4f1": 581,
    "c4f4": 582,
    "c4f7": 583,
    "c4g4": 584,
    "c4g8": 585,
    "c4h4": 586,
    "c5a3": 587,
    "c5a4": 588,
    "c5a5": 589,
    "c5a6": 590,
    "c5a7": 591,
    "c5b3": 592,
    "c5b4": 593,
    "c5b5": 594,
    "c5b6": 595,
    "c5b7": 596,
    "c5c1": 597,
    "c5c2": 598,
    "c5c3": 599,
    "c5c4": 600,
    "c5c6": 601,
    "c5c7": 602,
    "c5c8": 603,
    "c5d3": 604,
    "c5d4": 605,
    "c5d5": 606,
    "c5d6": 607,
    "c5d7": 608,
    "c5e3": 609,
    "c5e4": 610,
    "c5e5": 611,
    "c5e6": 612,
    "c5e7": 613,
    "c5f2": 614,
    "c5f5": 615,
    "c5f8": 616,
    "c5g1": 617,
    "c5g5": 618,
    "c5h5": 619,
    "c6a4": 620,
    "c6a5": 621,
    "c6a6": 622,
    "c6a7": 623,
    "c6a8": 624,
    "c6b4": 625,
    "c6b5": 626,
    "c6b6": 627,
    "c6b7": 628,
    "c6b8": 629,
    "c6c1": 630,
    "c6c2": 631,
    "c6c3": 632,
    "c6c4": 633,
    "c6c5": 634,
    "c6c7": 635,
    "c6c8": 636,
    "c6d4": 637,
    "c6d5": 638,
    "c6d6": 639,
    "c6d7": 640,
    "c6d8": 641,
    "c6e4": 642,
    "c6e5": 643,
    "c6e6": 644,
    "c6e7": 645,
    "c6e8": 646,
    "c6f3": 647,
    "c6f6": 648,
    "c6g2": 649,
    "c6g6": 650,
    "c6h1": 651,
    "c6h6": 652,
    "c7a5": 653,
    "c7a6": 654,
    "c7a7": 655,
    "c7a8": 656,
    "c7b5": 657,
    "c7b6": 658,
    "c7b7": 659,
    "c7b8": 660,
    "c7b8b": 661,
    "c7b8n": 662,
    "c7b8q": 663,
    "c7b8r": 664,
    "c7c1": 665,
    "c7c2": 666,
    "c7c3": 667,
    "c7c4": 668,
    "c7c5": 669,
    "c7c6": 670,
    "c7c8": 671,
    "c7c8b": 672,
    "c7c8n": 673,
    "c7c8q": 674,
    "c7c8r": 675,
    "c7d5": 676,
    "c7d6": 677,
    "c7d7": 678,
    "c7d8": 679,
    "c7d8b": 680,
    "c7d8n": 681,
    "c7d8q": 682,
    "c7d8r": 683,
    "c7e5": 684,
    "c7e6": 685,
    "c7e7": 686,
    "c7e8": 687,
    "c7f4": 688,
    "c7f7": 689,
    "c7g3": 690,
    "c7g7": 691,
    "c7h2": 692,
    "c7h7": 693,
    "c8a6": 694,
    "c8a7": 695,
    "c8a8": 696,
    "c8b6": 697,
    "c8b7": 698,
    "c8b8": 699,
    "c8c1": 700,
    "c8c2": 701,
    "c8c3": 702,
    "c8c4": 703,
    "c8c5": 704,
    "c8c6": 705,
    "c8c7": 706,
    "c8d6": 707,
    "c8d7": 708,
    "c8d8": 709,
    "c8e6": 710,
    "c8e7": 711,
    "c8e8": 712,
    "c8f5": 713,
    "c8f8": 714,
    "c8g4": 715,
    "c8g8": 716,
    "c8h3": 717,
    "c8h8": 718,
    "d1a1": 719,
    "d1a4": 720,
    "d1b1": 721,
    "d1b2": 722,
    "d1b3": 723,
    "d1c1": 724,
    "d1c2": 725,
    "d1c3": 726,
    "d1d2": 727,
    "d1d3": 728,
    "d1d4": 729,
    "d1d5": 730,
    "d1d6": 731,
    "d1d7": 732,
    "d1d8": 733,
    "d1e1": 734,
    "d1e2": 735,
    "d1e3": 736,
    "d1f1": 737,
    "d1f2": 738,
    "d1f3": 739,
    "d1g1": 740,
    "d1g4": 741,
    "d1h1": 742,
    "d1h5": 743,
    "d2a2": 744,
    "d2a5": 745,
    "d2b1": 746,
    "d2b2": 747,
    "d2b3": 748,
    "d2b4": 749,
    "d2c1": 750,
    "d2c1b": 751,
    "d2c1n": 752,
    "d2c1q": 753,
    "d2c1r": 754,
    "d2c2": 755,
    "d2c3": 756,
    "d2c4": 757,
    "d2d1": 758,
    "d2d1b": 759,
    "d2d1n": 760,
    "d2d1q": 761,
    "d2d1r": 762,
    "d2d3": 763,
    "d2d4": 764,
    "d2d5": 765,
    "d2d6": 766,
    "d2d7": 767,
    "d2d8": 768,
    "d2e1": 769,
    "d2e1b": 770,
    "d2e1n": 771,
    "d2e1q": 772,
    "d2e1r": 773,
    "d2e2": 774,
    "d2e3": 775,
    "d2e4": 776,
    "d2f1": 777,
    "d2f2": 778,
    "d2f3": 779,
    "d2f4": 780,
    "d2g2": 781,
    "d2g5": 782,
    "d2h2": 783,
    "d2h6": 784,
    "d3a3": 785,
    "d3a6": 786,
    "d3b1": 787,
    "d3b2": 788,
    "d3b3": 789,
    "d3b4": 790,
    "d3b5": 791,
    "d3c1": 792,
    "d3c2": 793,
    "d3c3": 794,
    "d3c4": 795,
    "d3c5": 796,
    "d3d1": 797,
    "d3d2": 798,
    "d3d4": 799,
    "d3d5": 800,
    "d3d6": 801,
    "d3d7": 802,
    "d3d8": 803,
    "d3e1": 804,
    "d3e2": 805,
    "d3e3": 806,
    "d3e4": 807,
    "d3e5": 808,
    "d3f1": 809,
    "d3f2": 810,
    "d3f3": 811,
    "d3f4": 812,
    "d3f5": 813,
    "d3g3": 814,
    "d3g6": 815,
    "d3h3": 816,
    "d3h7": 817,
    "d4a1": 818,
    "d4a4": 819,
    "d4a7": 820,
    "d4b2": 821,
    "d4b3": 822,
    "d4b4": 823,
    "d4b5": 824,
    "d4b6": 825,
    "d4c2": 826,
    "d4c3": 827,
    "d4c4": 828,
    "d4c5": 829,
    "d4c6": 830,
    "d4d1": 831,
    "d4d2": 832,
    "d4d3": 833,
    "d4d5": 834,
    "d4d6": 835,
    "d4d7": 836,
    "d4d8": 837,
    "d4e2": 838,
    "d4e3": 839,
    "d4e4": 840,
    "d4e5": 841,
    "d4e6": 842,
    "d4f2": 843,
    "d4f3": 844,
    "d4f4": 845,
    "d4f5": 846,
    "d4f6": 847,
    "d4g1": 848,
    "d4g4": 849,
    "d4g7": 850,
    "d4h4": 851,
    "d4h8": 852,
    "d5a2": 853,
    "d5a5": 854,
    "d5a8": 855,
    "d5b3": 856,
    "d5b4": 857,
    "d5b5": 858,
    "d5b6": 859,
    "d5b7": 860,
    "d5c3": 861,
    "d5c4": 862,
    "d5c5": 863,
    "d5c6": 864,
    "d5c7": 865,
    "d5d1": 866,
    "d5d2": 867,
    "d5d3": 868,
    "d5d4": 869,
    "d5d6": 870,
    "d5d7": 871,
    "d5d8": 872,
    "d5e3": 873,
    "d5e4": 874,
    "d5e5": 875,
    "d5e6": 876,
    "d5e7": 877,
    "d5f3": 878,
    "d5f4": 879,
    "d5f5": 880,
    "d5f6": 881,
    "d5f7": 882,
    "d5g2": 883,
    "d5g5": 884,
    "d5g8": 885,
    "d5h1": 886,
    "d5h5": 887,
    "d6a3": 888,
    "d6a6": 889,
    "d6b4": 890,
    "d6b5": 891,
    "d6b6": 892,
    "d6b7": 893,
    "d6b8": 894,
    "d6c4": 895,
    "d6c5": 896,
    "d6c6": 897,
    "d6c7": 898,
    "d6c8": 899,
    "d6d1": 900,
    "d6d2": 901,
    "d6d3": 902,
    "d6d4": 903,
    "d6d5": 904,
    "d6d7": 905,
    "d6d8": 906,
    "d6e4": 907,
    "d6e5": 908,
    "d6e6": 909,
    "d6e7": 910,
    "d6e8": 911,
    "d6f4": 912,
    "d6f5": 913,
    "d6f6": 914,
    "d6f7": 915,
    "d6f8": 916,
    "d6g3": 917,
    "d6g6": 918,
    "d6h2": 919,
    "d6h6": 920,
    "d7a4": 921,
    "d7a7": 922,
    "d7b5": 923,
    "d7b6": 924,
    "d7b7": 925,
    "d7b8": 926,
    "d7c5": 927,
    "d7c6": 928,
    "d7c7": 929,
    "d7c8": 930,
    "d7c8b": 931,
    "d7c8n": 932,
    "d7c8q": 933,
    "d7c8r": 934,
    "d7d1": 935,
    "d7d2": 936,
    "d7d3": 937,
    "d7d4": 938,
    "d7d5": 939,
    "d7d6": 940,
    "d7d8": 941,
    "d7d8b": 942,
    "d7d8n": 943,
    "d7d8q": 944,
    "d7d8r": 945,
    "d7e5": 946,
    "d7e6": 947,
    "d7e7": 948,
    "d7e8": 949,
    "d7e8b": 950,
    "d7e8n": 951,
    "d7e8q": 952,
    "d7e8r": 953,
    "d7f5": 954,
    "d7f6": 955,
    "d7f7": 956,
    "d7f8": 957,
    "d7g4": 958,
    "d7g7": 959,
    "d7h3": 960,
    "d7h7": 961,
    "d8a5": 962,
    "d8a8": 963,
    "d8b6": 964,
    "d8b7": 965,
    "d8b8": 966,
    "d8c6": 967,
    "d8c7": 968,
    "d8c8": 969,
    "d8d1": 970,
    "d8d2": 971,
    "d8d3": 972,
    "d8d4": 973,
    "d8d5": 974,
    "d8d6": 975,
    "d8d7": 976,
    "d8e6": 977,
    "d8e7": 978,
    "d8e8": 979,
    "d8f6": 980,
    "d8f7": 981,
    "d8f8": 982,
    "d8g5": 983,
    "d8g8": 984,
    "d8h4": 985,
    "d8h8": 986,
    "e1a1": 987,
    "e1a5": 988,
    "e1b1": 989,
    "e1b4": 990,
    "e1c1": 991,
    "e1c2": 992,
    "e1c3": 993,
    "e1d1": 994,
    "e1d2": 995,
    "e1d3": 996,
    "e1e2": 997,
    "e1e3": 998,
    "e1e4": 999,
    "e1e5": 1000,
    "e1e6": 1001,
    "e1e7": 1002,
    "e1e8": 1003,
    "e1f1": 1004,
    "e1f2": 1005,
    "e1f3": 1006,
    "e1g1": 1007,
    "e1g2": 1008,
    "e1g3": 1009,
    "e1h1": 1010,
    "e1h4": 1011,
    "e2a2": 1012,
    "e2a6": 1013,
    "e2b2": 1014,
    "e2b5": 1015,
    "e2c1": 1016,
    "e2c2": 1017,
    "e2c3": 1018,
    "e2c4": 1019,
    "e2d1": 1020,
    "e2d1b": 1021,
    "e2d1n": 1022,
    "e2d1q": 1023,
    "e2d1r": 1024,
    "e2d2": 1025,
    "e2d3": 1026,
    "e2d4": 1027,
    "e2e1": 1028,
    "e2e1b": 1029,
    "e2e1n": 1030,
    "e2e1q": 1031,
    "e2e1r": 1032,
    "e2e3": 1033,
    "e2e4": 1034,
    "e2e5": 1035,
    "e2e6": 1036,
    "e2e7": 1037,
    "e2e8": 1038,
    "e2f1": 1039,
    "e2f1b": 1040,
    "e2f1n": 1041,
    "e2f1q": 1042,
    "e2f1r": 1043,
    "e2f2": 1044,
    "e2f3": 1045,
    "e2f4": 1046,
    "e2g1": 1047,
    "e2g2": 1048,
    "e2g3": 1049,
    "e2g4": 1050,
    "e2h2": 1051,
    "e2h5": 1052,
    "e3a3": 1053,
    "e3a7": 1054,
    "e3b3": 1055,
    "e3b6": 1056,
    "e3c1": 1057,
    "e3c2": 1058,
    "e3c3": 1059,
    "e3c4": 1060,
    "e3c5": 1061,
    "e3d1": 1062,
    "e3d2": 1063,
    "e3d3": 1064,
    "e3d4": 1065,
    "e3d5": 1066,
    "e3e1": 1067,
    "e3e2": 1068,
    "e3e4": 1069,
    "e3e5": 1070,
    "e3e6": 1071,
    "e3e7": 1072,
    "e3e8": 1073,
    "e3f1": 1074,
    "e3f2": 1075,
    "e3f3": 1076,
    "e3f4": 1077,
    "e3f5": 1078,
    "e3g1": 1079,
    "e3g2": 1080,
    "e3g3": 1081,
    "e3g4": 1082,
    "e3g5": 1083,
    "e3h3": 1084,
    "e3h6": 1085,
    "e4a4": 1086,
    "e4a8": 1087,
    "e4b1": 1088,
    "e4b4": 1089,
    "e4b7": 1090,
    "e4c2": 1091,
    "e4c3": 1092,
    "e4c4": 1093,
    "e4c5": 1094,
    "e4c6": 1095,
    "e4d2": 1096,
    "e4d3": 1097,
    "e4d4": 1098,
    "e4d5": 1099,
    "e4d6": 1100,
    "e4e1": 1101,
    "e4e2": 1102,
    "e4e3": 1103,
    "e4e5": 1104,
    "e4e6": 1105,
    "e4e7": 1106,
    "e4e8": 1107,
    "e4f2": 1108,
    "e4f3": 1109,
    "e4f4": 1110,
    "e4f5": 1111,
    "e4f6": 1112,
    "e4g2": 1113,
    "e4g3": 1114,
    "e4g4": 1115,
    "e4g5": 1116,
    "e4g6": 1117,
    "e4h1": 1118,
    "e4h4": 1119,
    "e4h7": 1120,
    "e5a1": 1121,
    "e5a5": 1122,
    "e5b2": 1123,
    "e5b5": 1124,
    "e5b8": 1125,
    "e5c3": 1126,
    "e5c4": 1127,
    "e5c5": 1128,
    "e5c6": 1129,
    "e5c7": 1130,
    "e5d3": 1131,
    "e5d4": 1132,
    "e5d5": 1133,
    "e5d6": 1134,
    "e5d7": 1135,
    "e5e1": 1136,
    "e5e2": 1137,
    "e5e3": 1138,
    "e5e4": 1139,
    "e5e6": 1140,
    "e5e7": 1141,
    "e5e8": 1142,
    "e5f3": 1143,
    "e5f4": 1144,
    "e5f5": 1145,
    "e5f6": 1146,
    "e5f7": 1147,
    "e5g3": 1148,
    "e5g4": 1149,
    "e5g5": 1150,
    "e5g6": 1151,
    "e5g7": 1152,
    "e5h2": 1153,
    "e5h5": 1154,
    "e5h8": 1155,
    "e6a2": 1156,
    "e6a6": 1157,
    "e6b3": 1158,
    "e6b6": 1159,
    "e6c4": 1160,
    "e6c5": 1161,
    "e6c6": 1162,
    "e6c7": 1163,
    "e6c8": 1164,
    "e6d4": 1165,
    "e6d5": 1166,
    "e6d6": 1167,
    "e6d7": 1168,
    "e6d8": 1169,
    "e6e1": 1170,
    "e6e2": 1171,
    "e6e3": 1172,
    "e6e4": 1173,
    "e6e5": 1174,
    "e6e7": 1175,
    "e6e8": 1176,
    "e6f4": 1177,
    "e6f5": 1178,
    "e6f6": 1179,
    "e6f7": 1180,
    "e6f8": 1181,
    "e6g4": 1182,
    "e6g5": 1183,
    "e6g6": 1184,
    "e6g7": 1185,
    "e6g8": 1186,
    "e6h3": 1187,
    "e6h6": 1188,
    "e7a3": 1189,
    "e7a7": 1190,
    "e7b4": 1191,
    "e7b7": 1192,
    "e7c5": 1193,
    "e7c6": 1194,
    "e7c7": 1195,
    "e7c8": 1196,
    "e7d5": 1197,
    "e7d6": 1198,
    "e7d7": 1199,
    "e7d8": 1200,
    "e7d8b": 1201,
    "e7d8n": 1202,
    "e7d8q": 1203,
    "e7d8r": 1204,
    "e7e1": 1205,
    "e7e2": 1206,
    "e7e3": 1207,
    "e7e4": 1208,
    "e7e5": 1209,
    "e7e6": 1210,
    "e7e8": 1211,
    "e7e8b": 1212,
    "e7e8n": 1213,
    "e7e8q": 1214,
    "e7e8r": 1215,
    "e7f5": 1216,
    "e7f6": 1217,
    "e7f7": 1218,
    "e7f8": 1219,
    "e7f8b": 1220,
    "e7f8n": 1221,
    "e7f8q": 1222,
    "e7f8r": 1223,
    "e7g5": 1224,
    "e7g6": 1225,
    "e7g7": 1226,
    "e7g8": 1227,
    "e7h4": 1228,
    "e7h7": 1229,
    "e8a4": 1230,
    "e8a8": 1231,
    "e8b5": 1232,
    "e8b8": 1233,
    "e8c6": 1234,
    "e8c7": 1235,
    "e8c8": 1236,
    "e8d6": 1237,
    "e8d7": 1238,
    "e8d8": 1239,
    "e8e1": 1240,
    "e8e2": 1241,
    "e8e3": 1242,
    "e8e4": 1243,
    "e8e5": 1244,
    "e8e6": 1245,
    "e8e7": 1246,
    "e8f6": 1247,
    "e8f7": 1248,
    "e8f8": 1249,
    "e8g6": 1250,
    "e8g7": 1251,
    "e8g8": 1252,
    "e8h5": 1253,
    "e8h8": 1254,
    "f1a1": 1255,
    "f1a6": 1256,
    "f1b1": 1257,
    "f1b5": 1258,
    "f1c1": 1259,
    "f1c4": 1260,
    "f1d1": 1261,
    "f1d2": 1262,
    "f1d3": 1263,
    "f1e1": 1264,
    "f1e2": 1265,
    "f1e3": 1266,
    "f1f2": 1267,
    "f1f3": 1268,
    "f1f4": 1269,
    "f1f5": 1270,
    "f1f6": 1271,
    "f1f7": 1272,
    "f1f8": 1273,
    "f1g1": 1274,
    "f1g2": 1275,
    "f1g3": 1276,
    "f1h1": 1277,
    "f1h2": 1278,
    "f1h3": 1279,
    "f2a2": 1280,
    "f2a7": 1281,
    "f2b2": 1282,
    "f2b6": 1283,
    "f2c2": 1284,
    "f2c5": 1285,
    "f2d1": 1286,
    "f2d2": 1287,
    "f2d3": 1288,
    "f2d4": 1289,
    "f2e1": 1290,
    "f2e1b": 1291,
    "f2e1n": 1292,
    "f2e1q": 1293,
    "f2e1r": 1294,
    "f2e2": 1295,
    "f2e3": 1296,
    "f2e4": 1297,
    "f2f1": 1298,
    "f2f1b": 1299,
    "f2f1n": 1300,
    "f2f1q": 1301,
    "f2f1r": 1302,
    "f2f3": 1303,
    "f2f4": 1304,
    "f2f5": 1305,
    "f2f6": 1306,
    "f2f7": 1307,
    "f2f8": 1308,
    "f2g1": 1309,
    "f2g1b": 1310,
    "f2g1n": 1311,
    "f2g1q": 1312,
    "f2g1r": 1313,
    "f2g2": 1314,
    "f2g3": 1315,
    "f2g4": 1316,
    "f2h1": 1317,
    "f2h2": 1318,
    "f2h3": 1319,
    "f2h4": 1320,
    "f3a3": 1321,
    "f3a8": 1322,
    "f3b3": 1323,
    "f3b7": 1324,
    "f3c3": 1325,
    "f3c6": 1326,
    "f3d1": 1327,
    "f3d2": 1328,
    "f3d3": 1329,
    "f3d4": 1330,
    "f3d5": 1331,
    "f3e1": 1332,
    "f3e2": 1333,
    "f3e3": 1334,
    "f3e4": 1335,
    "f3e5": 1336,
    "f3f1": 1337,
    "f3f2": 1338,
    "f3f4": 1339,
    "f3f5": 1340,
    "f3f6": 1341,
    "f3f7": 1342,
    "f3f8": 1343,
    "f3g1": 1344,
    "f3g2": 1345,
    "f3g3": 1346,
    "f3g4": 1347,
    "f3g5": 1348,
    "f3h1": 1349,
    "f3h2": 1350,
    "f3h3": 1351,
    "f3h4": 1352,
    "f3h5": 1353,
    "f4a4": 1354,
    "f4b4": 1355,
    "f4b8": 1356,
    "f4c1": 1357,
    "f4c4": 1358,
    "f4c7": 1359,
    "f4d2": 1360,
    "f4d3": 1361,
    "f4d4": 1362,
    "f4d5": 1363,
    "f4d6": 1364,
    "f4e2": 1365,
    "f4e3": 1366,
    "f4e4": 1367,
    "f4e5": 1368,
    "f4e6": 1369,
    "f4f1": 1370,
    "f4f2": 1371,
    "f4f3": 1372,
    "f4f5": 1373,
    "f4f6": 1374,
    "f4f7": 1375,
    "f4f8": 1376,
    "f4g2": 1377,
    "f4g3": 1378,
    "f4g4": 1379,
    "f4g5": 1380,
    "f4g6": 1381,
    "f4h2": 1382,
    "f4h3": 1383,
    "f4h4": 1384,
    "f4h5": 1385,
    "f4h6": 1386,
    "f5a5": 1387,
    "f5b1": 1388,
    "f5b5": 1389,
    "f5c2": 1390,
    "f5c5": 1391,
    "f5c8": 1392,
    "f5d3": 1393,
    "f5d4": 1394,
    "f5d5": 1395,
    "f5d6": 1396,
    "f5d7": 1397,
    "f5e3": 1398,
    "f5e4": 1399,
    "f5e5": 1400,
    "f5e6": 1401,
    "f5e7": 1402,
    "f5f1": 1403,
    "f5f2": 1404,
    "f5f3": 1405,
    "f5f4": 1406,
    "f5f6": 1407,
    "f5f7": 1408,
    "f5f8": 1409,
    "f5g3": 1410,
    "f5g4": 1411,
    "f5g5": 1412,
    "f5g6": 1413,
    "f5g7": 1414,
    "f5h3": 1415,
    "f5h4": 1416,
    "f5h5": 1417,
    "f5h6": 1418,
    "f5h7": 1419,
    "f6a1": 1420,
    "f6a6": 1421,
    "f6b2": 1422,
    "f6b6": 1423,
    "f6c3": 1424,
    "f6c6": 1425,
    "f6d4": 1426,
    "f6d5": 1427,
    "f6d6": 1428,
    "f6d7": 1429,
    "f6d8": 1430,
    "f6e4": 1431,
    "f6e5": 1432,
    "f6e6": 1433,
    "f6e7": 1434,
    "f6e8": 1435,
    "f6f1": 1436,
    "f6f2": 1437,
    "f6f3": 1438,
    "f6f4": 1439,
    "f6f5": 1440,
    "f6f7": 1441,
    "f6f8": 1442,
    "f6g4": 1443,
    "f6g5": 1444,
    "f6g6": 1445,
    "f6g7": 1446,
    "f6g8": 1447,
    "f6h4": 1448,
    "f6h5": 1449,
    "f6h6": 1450,
    "f6h7": 1451,
    "f6h8": 1452,
    "f7a2": 1453,
    "f7a7": 1454,
    "f7b3": 1455,
    "f7b7": 1456,
    "f7c4": 1457,
    "f7c7": 1458,
    "f7d5": 1459,
    "f7d6": 1460,
    "f7d7": 1461,
    "f7d8": 1462,
    "f7e5": 1463,
    "f7e6": 1464,
    "f7e7": 1465,
    "f7e8": 1466,
    "f7e8b": 1467,
    "f7e8n": 1468,
    "f7e8q": 1469,
    "f7e8r": 1470,
    "f7f1": 1471,
    "f7f2": 1472,
    "f7f3": 1473,
    "f7f4": 1474,
    "f7f5": 1475,
    "f7f6": 1476,
    "f7f8": 1477,
    "f7f8b": 1478,
    "f7f8n": 1479,
    "f7f8q": 1480,
    "f7f8r": 1481,
    "f7g5": 1482,
    "f7g6": 1483,
    "f7g7": 1484,
    "f7g8": 1485,
    "f7g8b": 1486,
    "f7g8n": 1487,
    "f7g8q": 1488,
    "f7g8r": 1489,
    "f7h5": 1490,
    "f7h6": 1491,
    "f7h7": 1492,
    "f7h8": 1493,
    "f8a3": 1494,
    "f8a8": 1495,
    "f8b4": 1496,
    "f8b8": 1497,
    "f8c5": 1498,
    "f8c8": 1499,
    "f8d6": 1500,
    "f8d7": 1501,
    "f8d8": 1502,
    "f8e6": 1503,
    "f8e7": 1504,
    "f8e8": 1505,
    "f8f1": 1506,
    "f8f2": 1507,
    "f8f3": 1508,
    "f8f4": 1509,
    "f8f5": 1510,
    "f8f6": 1511,
    "f8f7": 1512,
    "f8g6": 1513,
    "f8g7": 1514,
    "f8g8": 1515,
    "f8h6": 1516,
    "f8h7": 1517,
    "f8h8": 1518,
    "g1a1": 1519,
    "g1a7": 1520,
    "g1b1": 1521,
    "g1b6": 1522,
    "g1c1": 1523,
    "g1c5": 1524,
    "g1d1": 1525,
    "g1d4": 1526,
    "g1e1": 1527,
    "g1e2": 1528,
    "g1e3": 1529,
    "g1f1": 1530,
    "g1f2": 1531,
    "g1f3": 1532,
    "g1g2": 1533,
    "g1g3": 1534,
    "g1g4": 1535,
    "g1g5": 1536,
    "g1g6": 1537,
    "g1g7": 1538,
    "g1g8": 1539,
    "g1h1": 1540,
    "g1h2": 1541,
    "g1h3": 1542,
    "g2a2": 1543,
    "g2a8": 1544,
    "g2b2": 1545,
    "g2b7": 1546,
    "g2c2": 1547,
    "g2c6": 1548,
    "g2d2": 1549,
    "g2d5": 1550,
    "g2e1": 1551,
    "g2e2": 1552,
    "g2e3": 1553,
    "g2e4": 1554,
    "g2f1": 1555,
    "g2f1b": 1556,
    "g2f1n": 1557,
    "g2f1q": 1558,
    "g2f1r": 1559,
    "g2f2": 1560,
    "g2f3": 1561,
    "g2f4": 1562,
    "g2g1": 1563,
    "g2g1b": 1564,
    "g2g1n": 1565,
    "g2g1q": 1566,
    "g2g1r": 1567,
    "g2g3": 1568,
    "g2g4": 1569,
    "g2g5": 1570,
    "g2g6": 1571,
    "g2g7": 1572,
    "g2g8": 1573,
    "g2h1": 1574,
    "g2h1b": 1575,
    "g2h1n": 1576,
    "g2h1q": 1577,
    "g2h1r": 1578,
    "g2h2": 1579,
    "g2h3": 1580,
    "g2h4": 1581,
    "g3a3": 1582,
    "g3b3": 1583,
    "g3b8": 1584,
    "g3c3": 1585,
    "g3c7": 1586,
    "g3d3": 1587,
    "g3d6": 1588,
    "g3e1": 1589,
    "g3e2": 1590,
    "g3e3": 1591,
    "g3e4": 1592,
    "g3e5": 1593,
    "g3f1": 1594,
    "g3f2": 1595,
    "g3f3": 1596,
    "g3f4": 1597,
    "g3f5": 1598,
    "g3g1": 1599,
    "g3g2": 1600,
    "g3g4": 1601,
    "g3g5": 1602,
    "g3g6": 1603,
    "g3g7": 1604,
    "g3g8": 1605,
    "g3h1": 1606,
    "g3h2": 1607,
    "g3h3": 1608,
    "g3h4": 1609,
    "g3h5": 1610,
    "g4a4": 1611,
    "g4b4": 1612,
    "g4c4": 1613,
    "g4c8": 1614,
    "g4d1": 1615,
    "g4d4": 1616,
    "g4d7": 1617,
    "g4e2": 1618,
    "g4e3": 1619,
    "g4e4": 1620,
    "g4e5": 1621,
    "g4e6": 1622,
    "g4f2": 1623,
    "g4f3": 1624,
    "g4f4": 1625,
    "g4f5": 1626,
    "g4f6": 1627,
    "g4g1": 1628,
    "g4g2": 1629,
    "g4g3": 1630,
    "g4g5": 1631,
    "g4g6": 1632,
    "g4g7": 1633,
    "g4g8": 1634,
    "g4h2": 1635,
    "g4h3": 1636,
    "g4h4": 1637,
    "g4h5": 1638,
    "g4h6": 1639,
    "g5a5": 1640,
    "g5b5": 1641,
    "g5c1": 1642,
    "g5c5": 1643,
    "g5d2": 1644,
    "g5d5": 1645,
    "g5d8": 1646,
    "g5e3": 1647,
    "g5e4": 1648,
    "g5e5": 1649,
    "g5e6": 1650,
    "g5e7": 1651,
    "g5f3": 1652,
    "g5f4": 1653,
    "g5f5": 1654,
    "g5f6": 1655,
    "g5f7": 1656,
    "g5g1": 1657,
    "g5g2": 1658,
    "g5g3": 1659,
    "g5g4": 1660,
    "g5g6": 1661,
    "g5g7": 1662,
    "g5g8": 1663,
    "g5h3": 1664,
    "g5h4": 1665,
    "g5h5": 1666,
    "g5h6": 1667,
    "g5h7": 1668,
    "g6a6": 1669,
    "g6b1": 1670,
    "g6b6": 1671,
    "g6c2": 1672,
    "g6c6": 1673,
    "g6d3": 1674,
    "g6d6": 1675,
    "g6e4": 1676,
    "g6e5": 1677,
    "g6e6": 1678,
    "g6e7": 1679,
    "g6e8": 1680,
    "g6f4": 1681,
    "g6f5": 1682,
    "g6f6": 1683,
    "g6f7": 1684,
    "g6f8": 1685,
    "g6g1": 1686,
    "g6g2": 1687,
    "g6g3": 1688,
    "g6g4": 1689,
    "g6g5": 1690,
    "g6g7": 1691,
    "g6g8": 1692,
    "g6h4": 1693,
    "g6h5": 1694,
    "g6h6": 1695,
    "g6h7": 1696,
    "g6h8": 1697,
    "g7a1": 1698,
    "g7a7": 1699,
    "g7b2": 1700,
    "g7b7": 1701,
    "g7c3": 1702,
    "g7c7": 1703,
    "g7d4": 1704,
    "g7d7": 1705,
    "g7e5": 1706,
    "g7e6": 1707,
    "g7e7": 1708,
    "g7e8": 1709,
    "g7f5": 1710,
    "g7f6": 1711,
    "g7f7": 1712,
    "g7f8": 1713,
    "g7f8b": 1714,
    "g7f8n": 1715,
    "g7f8q": 1716,
    "g7f8r": 1717,
    "g7g1": 1718,
    "g7g2": 1719,
    "g7g3": 1720,
    "g7g4": 1721,
    "g7g5": 1722,
    "g7g6": 1723,
    "g7g8": 1724,
    "g7g8b": 1725,
    "g7g8n": 1726,
    "g7g8q": 1727,
    "g7g8r": 1728,
    "g7h5": 1729,
    "g7h6": 1730,
    "g7h7": 1731,
    "g7h8": 1732,
    "g7h8b": 1733,
    "g7h8n": 1734,
    "g7h8q": 1735,
    "g7h8r": 1736,
    "g8a2": 1737,
    "g8a8": 1738,
    "g8b3": 1739,
    "g8b8": 1740,
    "g8c4": 1741,
    "g8c8": 1742,
    "g8d5": 1743,
    "g8d8": 1744,
    "g8e6": 1745,
    "g8e7": 1746,
    "g8e8": 1747,
    "g8f6": 1748,
    "g8f7": 1749,
    "g8f8": 1750,
    "g8g1": 1751,
    "g8g2": 1752,
    "g8g3": 1753,
    "g8g4": 1754,
    "g8g5": 1755,
    "g8g6": 1756,
    "g8g7": 1757,
    "g8h6": 1758,
    "g8h7": 1759,
    "g8h8": 1760,
    "h1a1": 1761,
    "h1a8": 1762,
    "h1b1": 1763,
    "h1b7": 1764,
    "h1c1": 1765,
    "h1c6": 1766,
    "h1d1": 1767,
    "h1d5": 1768,
    "h1e1": 1769,
    "h1e4": 1770,
    "h1f1": 1771,
    "h1f2": 1772,
    "h1f3": 1773,
    "h1g1": 1774,
    "h1g2": 1775,
    "h1g3": 1776,
    "h1h2": 1777,
    "h1h3": 1778,
    "h1h4": 1779,
    "h1h5": 1780,
    "h1h6": 1781,
    "h1h7": 1782,
    "h1h8": 1783,
    "h2a2": 1784,
    "h2b2": 1785,
    "h2b8": 1786,
    "h2c2": 1787,
    "h2c7": 1788,
    "h2d2": 1789,
    "h2d6": 1790,
    "h2e2": 1791,
    "h2e5": 1792,
    "h2f1": 1793,
    "h2f2": 1794,
    "h2f3": 1795,
    "h2f4": 1796,
    "h2g1": 1797,
    "h2g1b": 1798,
    "h2g1n": 1799,
    "h2g1q": 1800,
    "h2g1r": 1801,
    "h2g2": 1802,
    "h2g3": 1803,
    "h2g4": 1804,
    "h2h1": 1805,
    "h2h1b": 1806,
    "h2h1n": 1807,
    "h2h1q": 1808,
    "h2h1r": 1809,
    "h2h3": 1810,
    "h2h4": 1811,
    "h2h5": 1812,
    "h2h6": 1813,
    "h2h7": 1814,
    "h2h8": 1815,
    "h3a3": 1816,
    "h3b3": 1817,
    "h3c3": 1818,
    "h3c8": 1819,
    "h3d3": 1820,
    "h3d7": 1821,
    "h3e3": 1822,
    "h3e6": 1823,
    "h3f1": 1824,
    "h3f2": 1825,
    "h3f3": 1826,
    "h3f4": 1827,
    "h3f5": 1828,
    "h3g1": 1829,
    "h3g2": 1830,
    "h3g3": 1831,
    "h3g4": 1832,
    "h3g5": 1833,
    "h3h1": 1834,
    "h3h2": 1835,
    "h3h4": 1836,
    "h3h5": 1837,
    "h3h6": 1838,
    "h3h7": 1839,
    "h3h8": 1840,
    "h4a4": 1841,
    "h4b4": 1842,
    "h4c4": 1843,
    "h4d4": 1844,
    "h4d8": 1845,
    "h4e1": 1846,
    "h4e4": 1847,
    "h4e7": 1848,
    "h4f2": 1849,
    "h4f3": 1850,
    "h4f4": 1851,
    "h4f5": 1852,
    "h4f6": 1853,
    "h4g2": 1854,
    "h4g3": 1855,
    "h4g4": 1856,
    "h4g5": 1857,
    "h4g6": 1858,
    "h4h1": 1859,
    "h4h2": 1860,
    "h4h3": 1861,
    "h4h5": 1862,
    "h4h6": 1863,
    "h4h7": 1864,
    "h4h8": 1865,
    "h5a5": 1866,
    "h5b5": 1867,
    "h5c5": 1868,
    "h5d1": 1869,
    "h5d5": 1870,
    "h5e2": 1871,
    "h5e5": 1872,
    "h5e8": 1873,
    "h5f3": 1874,
    "h5f4": 1875,
    "h5f5": 1876,
    "h5f6": 1877,
    "h5f7": 1878,
    "h5g3": 1879,
    "h5g4": 1880,
    "h5g5": 1881,
    "h5g6": 1882,
    "h5g7": 1883,
    "h5h1": 1884,
    "h5h2": 1885,
    "h5h3": 1886,
    "h5h4": 1887,
    "h5h6": 1888,
    "h5h7": 1889,
    "h5h8": 1890,
    "h6a6": 1891,
    "h6b6": 1892,
    "h6c1": 1893,
    "h6c6": 1894,
    "h6d2": 1895,
    "h6d6": 1896,
    "h6e3": 1897,
    "h6e6": 1898,
    "h6f4": 1899,
    "h6f5": 1900,
    "h6f6": 1901,
    "h6f7": 1902,
    "h6f8": 1903,
    "h6g4": 1904,
    "h6g5": 1905,
    "h6g6": 1906,
    "h6g7": 1907,
    "h6g8": 1908,
    "h6h1": 1909,
    "h6h2": 1910,
    "h6h3": 1911,
    "h6h4": 1912,
    "h6h5": 1913,
    "h6h7": 1914,
    "h6h8": 1915,
    "h7a7": 1916,
    "h7b1": 1917,
    "h7b7": 1918,
    "h7c2": 1919,
    "h7c7": 1920,
    "h7d3": 1921,
    "h7d7": 1922,
    "h7e4": 1923,
    "h7e7": 1924,
    "h7f5": 1925,
    "h7f6": 1926,
    "h7f7": 1927,
    "h7f8": 1928,
    "h7g5": 1929,
    "h7g6": 1930,
    "h7g7": 1931,
    "h7g8": 1932,
    "h7g8b": 1933,
    "h7g8n": 1934,
    "h7g8q": 1935,
    "h7g8r": 1936,
    "h7h1": 1937,
    "h7h2": 1938,
    "h7h3": 1939,
    "h7h4": 1940,
    "h7h5": 1941,
    "h7h6": 1942,
    "h7h8": 1943,
    "h7h8b": 1944,
    "h7h8n": 1945,
    "h7h8q": 1946,
    "h7h8r": 1947,
    "h8a1": 1948,
    "h8a8": 1949,
    "h8b2": 1950,
    "h8b8": 1951,
    "h8c3": 1952,
    "h8c8": 1953,
    "h8d4": 1954,
    "h8d8": 1955,
    "h8e5": 1956,
    "h8e8": 1957,
    "h8f6": 1958,
    "h8f7": 1959,
    "h8f8": 1960,
    "h8g6": 1961,
    "h8g7": 1962,
    "h8g8": 1963,
    "h8h1": 1964,
    "h8h2": 1965,
    "h8h3": 1966,
    "h8h4": 1967,
    "h8h5": 1968,
    "h8h6": 1969,
    "h8h7": 1970
  },
  "vocab_size": 1971,
  "n_special": 3,
  "n_moves": 1968,
  "pad_id": 0,
  "bos_id": 1,
  "eos_id": 2
}
vocab.py140 lines

Builds the vocabulary by geometric enumeration of every possible move. Produces the 1971 tokens.

"""Construction du vocabulaire : un token = un coup UCI.

C'est la décision de conception la plus structurante du projet, alors elle
mérite d'être expliquée.

Un modèle de langage classique découpe le texte en fragments de mots (BPE).
Ici on fait autrement : le vocabulaire est l'ensemble *fini et connu d'avance*
de tous les coups jouables aux échecs, écrits en notation UCI — case de départ
suivie de la case d'arrivée, plus éventuellement la pièce de promotion :
`e2e4`, `g1f3`, `e7e8q`. Un coup = un token, indivisible.

Pourquoi c'est mieux qu'un BPE ici : avec un BPE, `e2e4` pourrait se découper
en `e2` + `e4`, ou pire en `e` + `2e` + `4`, et le modèle devrait dépenser de
la capacité à réapprendre que ces fragments forment une unité. Avec un token
par coup, une partie de 80 demi-coups fait exactement 80 tokens, la fenêtre de
contexte se raisonne en coups plutôt qu'en caractères, et le masquage des
coups illégaux (phase 4) devient un simple masque booléen sur le vocabulaire.

Comment on énumère les coups possibles : plutôt que d'essayer de deviner quels
coups sont légaux (ça dépend de la position), on énumère tous les déplacements
*géométriquement* possibles sur un échiquier vide. Toute pièce se déplace soit
comme une dame (lignes, colonnes, diagonales), soit comme un cavalier — le roi,
la tour, le fou et le pion ne font que des sous-ensembles des déplacements de
la dame. On ajoute ensuite les promotions. Le vocabulaire obtenu est un
sur-ensemble strict des coups légaux, ce qui est exactement ce qu'on veut :
il ne manquera jamais un coup, et le modèle apprendra tout seul que certains
n'apparaissent jamais.
"""

import json

FILES = "abcdefgh"
RANKS = "12345678"

SPECIAL_TOKENS = ["<pad>", "<bos>", "<eos>"]

# Déplacements du cavalier, en (delta_colonne, delta_ligne)
KNIGHT_DELTAS = [(1, 2), (2, 1), (2, -1), (1, -2),
                 (-1, -2), (-2, -1), (-2, 1), (-1, 2)]

# Directions de la dame : les 8 rayons (horizontal, vertical, diagonal)
QUEEN_DIRECTIONS = [(0, 1), (1, 1), (1, 0), (1, -1),
                    (0, -1), (-1, -1), (-1, 0), (-1, 1)]

PROMOTION_PIECES = "qrbn"


def square_name(file_idx: int, rank_idx: int) -> str:
    """(0, 0) -> 'a1', (7, 7) -> 'h8'."""
    return FILES[file_idx] + RANKS[rank_idx]


def build_move_list() -> list[str]:
    """Énumère tous les coups UCI géométriquement possibles.

    Renvoie une liste triée, donc déterministe : le même vocabulaire sera
    reconstruit à l'identique sur n'importe quelle machine, ce qui évite les
    désastres silencieux du type « le checkpoint a été entraîné avec un autre
    ordre de tokens ».
    """
    moves = set()

    for from_file in range(8):
        for from_rank in range(8):
            origin = square_name(from_file, from_rank)

            # Déplacements de dame : on parcourt chaque rayon jusqu'au bord
            for d_file, d_rank in QUEEN_DIRECTIONS:
                for distance in range(1, 8):
                    to_file = from_file + d_file * distance
                    to_rank = from_rank + d_rank * distance
                    if not (0 <= to_file < 8 and 0 <= to_rank < 8):
                        break
                    moves.add(origin + square_name(to_file, to_rank))

            # Déplacements de cavalier
            for d_file, d_rank in KNIGHT_DELTAS:
                to_file = from_file + d_file
                to_rank = from_rank + d_rank
                if 0 <= to_file < 8 and 0 <= to_rank < 8:
                    moves.add(origin + square_name(to_file, to_rank))

    # Promotions. Un pion blanc promeut en allant de la 7e à la 8e rangée,
    # un pion noir de la 2e à la 1re. Il peut avancer tout droit ou capturer
    # en diagonale, d'où le décalage de colonne dans {-1, 0, +1}.
    for from_rank, to_rank in ((6, 7), (1, 0)):
        for from_file in range(8):
            for d_file in (-1, 0, 1):
                to_file = from_file + d_file
                if not 0 <= to_file < 8:
                    continue
                base = (square_name(from_file, from_rank)
                        + square_name(to_file, to_rank))
                for piece in PROMOTION_PIECES:
                    moves.add(base + piece)

    return sorted(moves)


def build_vocab() -> dict:
    """Construit le dictionnaire complet : tokens spéciaux puis coups."""
    move_list = build_move_list()
    itos = SPECIAL_TOKENS + move_list
    stoi = {token: idx for idx, token in enumerate(itos)}

    assert len(itos) == len(stoi), "collision dans le vocabulaire"
    # Le vocabulaire doit tenir sur 16 bits, puisqu'on encode le corpus en
    # uint16 pour diviser par deux la taille du fichier .bin.
    assert len(itos) < 2 ** 16, "vocabulaire trop grand pour un uint16"

    return {
        "itos": itos,
        "stoi": stoi,
        "vocab_size": len(itos),
        "n_special": len(SPECIAL_TOKENS),
        "n_moves": len(move_list),
        "pad_id": stoi["<pad>"],
        "bos_id": stoi["<bos>"],
        "eos_id": stoi["<eos>"],
    }


def save_vocab(path: str) -> dict:
    vocab = build_vocab()
    with open(path, "w") as f:
        json.dump(vocab, f, indent=2)
    return vocab


if __name__ == "__main__":
    import sys

    out = sys.argv[1] if len(sys.argv) > 1 else "data/vocab.json"
    v = save_vocab(out)
    print(f"Vocabulaire écrit dans {out}")
    print(f"  taille totale      : {v['vocab_size']}")
    print(f"  tokens spéciaux    : {v['n_special']} ({', '.join(SPECIAL_TOKENS)})")
    print(f"  coups UCI distincts: {v['n_moves']}")
    print(f"  exemples           : {', '.join(v['itos'][3:9])} ... "
          f"{', '.join(v['itos'][-4:])}")
prepare_data.py639 lines

Downloads, measures, filters, converts SAN to UCI and encodes to binary. Three subcommands: benchmark, parse, encode.

"""Phase 1 — Du dump Lichess brut à un corpus de parties en notation UCI.

Le dump mensuel de Lichess est un fichier PGN compressé en zstd d'environ 29 Go
(soit ~200 Go une fois décompressé, plus de 100 millions de parties). On veut
en extraire un sous-ensemble propre, écrit sous une forme que le modèle pourra
avaler directement : une partie par ligne, les coups en UCI séparés par des
espaces.

Deux opérations, de coûts très différents :

1. **Filtrer sur les en-têtes** — bon marché. Ce sont des comparaisons de
   chaînes sur les métadonnées de la partie (Elo des joueurs, cadence, cause de
   fin de partie).
2. **Convertir SAN vers UCI** — cher. Lichess stocke les coups en notation
   algébrique abrégée (`Nf3`), qui est *contextuelle* : « Nf3 » ne dit pas d'où
   vient le cavalier, il faut connaître la position pour le savoir. La seule
   façon de traduire en UCI (`g1f3`), qui est explicite, est de rejouer la
   partie coup par coup sur un échiquier. C'est ce qui coûte, et c'est
   incompressible.

D'où l'architecture : le processus principal ne fait que décompresser et
découper le flux en blocs de parties, puis distribue ces blocs à des processus
ouvriers qui font le filtrage *et* le rejeu. La décompression zstd est rapide
et reste séquentielle ; le rejeu, lui, est parallélisé sur les cœurs
disponibles.

Usage :
    python prepare_data.py benchmark --games 50000
    python prepare_data.py parse --workers 10 --out data/games_uci.txt
    python prepare_data.py encode --games data/games_uci.txt
"""

import argparse
import io
import json
import multiprocessing as mp
import os
import re
import sys
import time
from collections import Counter

import numpy as np
import zstandard

# --------------------------------------------------------------------------
# Critères de filtrage — ce sont les décisions figées du projet
# --------------------------------------------------------------------------

ELO_MIN = 1800
ELO_MAX = 2600
MIN_PLIES = 20          # demi-coups : en dessous, la partie n'a rien à apprendre
MAX_PLIES = 300         # au-dessus, c'est une anomalie (et ça dépasse le contexte)
EXCLUDED_SPEEDS = ("Bullet", "UltraBullet")
REQUIRED_TERMINATION = "Normal"

# --------------------------------------------------------------------------
# Extraction du texte des coups
# --------------------------------------------------------------------------

# Les commentaires Lichess : { [%eval 0.17] [%clk 0:03:00] }
COMMENT_RE = re.compile(r"\{[^}]*\}")
# Les numéros de coup : "1." pour les blancs, "1..." pour les noirs
MOVENUM_RE = re.compile(r"\d+\.+")
# Les résultats en fin de ligne
RESULT_TOKENS = {"1-0", "0-1", "1/2-1/2", "*"}
# Les annotations d'appréciation collées au coup : Nf3!? -> Nf3
ANNOTATION_RE = re.compile(r"[?!]+$")
# Les NAG numériques : $1, $18...
NAG_RE = re.compile(r"^\$\d+$")

HEADER_RE = re.compile(r'^\[(\w+)\s+"(.*)"\]$')


def extract_san_moves(movetext: str) -> list[str]:
    """Transforme le texte PGN des coups en une liste de coups SAN nus."""
    movetext = COMMENT_RE.sub(" ", movetext)
    movetext = MOVENUM_RE.sub(" ", movetext)
    moves = []
    for token in movetext.split():
        if token in RESULT_TOKENS or NAG_RE.match(token):
            continue
        token = ANNOTATION_RE.sub("", token)
        if token:
            moves.append(token)
    return moves


# --------------------------------------------------------------------------
# Traitement d'une partie (exécuté dans les processus ouvriers)
# --------------------------------------------------------------------------

# Codes de rejet, pour pouvoir expliquer le taux de rétention plutôt que de le
# constater. Chaque partie écartée l'est pour exactement une raison, la
# première rencontrée dans l'ordre ci-dessous.
REJECT_REASONS = [
    "cadence_exclue",       # bullet ou ultrabullet
    "elo_hors_bornes",      # au moins un joueur hors [1800, 2600]
    "elo_absent",           # en-tête WhiteElo/BlackElo manquant ou non numérique
    "terminaison",          # Termination != Normal
    "position_initiale",    # partie démarrée depuis une position custom (FEN)
    "trop_courte",          # < 20 demi-coups
    "trop_longue",          # > 300 demi-coups
    "san_invalide",         # coup impossible à rejouer (PGN corrompu)
]


def process_game_block(block: str, board_cls):
    """Filtre puis convertit une partie. Renvoie (uci_string | None, raison).

    `board_cls` est passé en argument plutôt qu'importé au niveau module pour
    que l'import de python-chess n'ait lieu qu'une fois par ouvrier.
    """
    headers = {}
    movetext_lines = []
    in_movetext = False

    for line in block.split("\n"):
        line = line.strip()
        if not line:
            if headers:
                in_movetext = True
            continue
        if not in_movetext and line.startswith("["):
            m = HEADER_RE.match(line)
            if m:
                headers[m.group(1)] = m.group(2)
        else:
            in_movetext = True
            movetext_lines.append(line)

    # --- Filtres bon marché, du plus discriminant au moins discriminant ---

    event = headers.get("Event", "")
    if any(speed in event for speed in EXCLUDED_SPEEDS):
        return None, "cadence_exclue"

    try:
        white_elo = int(headers["WhiteElo"])
        black_elo = int(headers["BlackElo"])
    except (KeyError, ValueError):
        return None, "elo_absent"

    if not (ELO_MIN <= white_elo <= ELO_MAX and ELO_MIN <= black_elo <= ELO_MAX):
        return None, "elo_hors_bornes"

    if headers.get("Termination") != REQUIRED_TERMINATION:
        return None, "terminaison"

    # Une partie démarrée depuis une position arbitraire fausserait
    # l'apprentissage : le modèle ne voit pas la position, seulement les coups,
    # donc il croirait que ces coups suivent le début de partie standard.
    if "FEN" in headers:
        return None, "position_initiale"

    san_moves = extract_san_moves(" ".join(movetext_lines))

    if len(san_moves) < MIN_PLIES:
        return None, "trop_courte"
    if len(san_moves) > MAX_PLIES:
        return None, "trop_longue"

    # --- Filtre cher : le rejeu ---

    board = board_cls()
    uci_moves = []
    try:
        for san in san_moves:
            move = board.parse_san(san)
            uci_moves.append(move.uci())
            board.push(move)
    except Exception:
        return None, "san_invalide"

    return " ".join(uci_moves), None


# --------------------------------------------------------------------------
# Ouvriers
# --------------------------------------------------------------------------

_BOARD_CLS = None


def _worker_init():
    """Importe python-chess une seule fois par processus ouvrier."""
    global _BOARD_CLS
    import chess
    _BOARD_CLS = chess.Board


def _worker_process_batch(blocks: list[str]):
    """Traite un lot de parties.

    Renvoie (lignes_uci, compteur_de_rejets, nombre_de_parties_lues). Le
    troisième élément peut sembler redondant avec la taille du lot, mais le
    dernier lot d'un flux est presque toujours incomplet : sans ce compteur,
    le taux de rétention publié serait faux de quelques dixièmes de pourcent.
    """
    kept = []
    rejected = Counter()
    for block in blocks:
        uci, reason = process_game_block(block, _BOARD_CLS)
        if uci is None:
            rejected[reason] += 1
        else:
            kept.append(uci)
    return kept, rejected, len(blocks)


# --------------------------------------------------------------------------
# Lecture du flux compressé
# --------------------------------------------------------------------------

def iter_game_blocks(path: str, read_size: int = 1 << 24):
    """Parcourt le .pgn.zst et produit un bloc de texte par partie.

    On ne décompresse jamais le fichier sur disque : `stream_reader` nous donne
    un flux, qu'on lit par tranches de 16 Mo. C'est ce qui permet de traiter
    200 Go de PGN avec 29 Go sur le disque et quelques centaines de Mo de RAM.

    Le découpage se fait sur la balise `[Event `, qui ouvre chaque partie.
    """
    dctx = zstandard.ZstdDecompressor()
    with open(path, "rb") as fh:
        with dctx.stream_reader(fh, read_size=read_size) as reader:
            text_stream = io.TextIOWrapper(reader, encoding="utf-8",
                                           errors="replace")
            current = []
            for line in text_stream:
                if line.startswith("[Event ") and current:
                    yield "".join(current)
                    current = [line]
                else:
                    current.append(line)
            if current:
                yield "".join(current)


def batched(iterable, size):
    batch = []
    for item in iterable:
        batch.append(item)
        if len(batch) >= size:
            yield batch
            batch = []
    if batch:
        yield batch


# --------------------------------------------------------------------------
# Commande : benchmark
# --------------------------------------------------------------------------

def cmd_benchmark(args):
    """Mesure le débit de parsing, en séquentiel et en parallèle.

    L'objectif est de répondre à une question concrète : combien de temps pour
    atteindre ~800 M de tokens de coups ? Si la réponse dépasse 8 heures en
    séquentiel, on parallélise.
    """
    import chess

    print(f"Lecture de {args.games} parties depuis {args.input}\n")

    # --- 1. Coût de la lecture seule (décompression + découpage) ---
    t0 = time.perf_counter()
    blocks = []
    for i, block in enumerate(iter_game_blocks(args.input)):
        blocks.append(block)
        if i + 1 >= args.games:
            break
    t_read = time.perf_counter() - t0
    print(f"[1] Décompression + découpage : {t_read:.1f} s "
          f"({len(blocks)/t_read:,.0f} parties/s)")

    # --- 2. Coût du traitement complet, un seul cœur ---
    t0 = time.perf_counter()
    kept_seq = []
    rejected = Counter()
    for block in blocks:
        uci, reason = process_game_block(block, chess.Board)
        if uci is None:
            rejected[reason] += 1
        else:
            kept_seq.append(uci)
    t_seq = time.perf_counter() - t0
    rate_seq = len(blocks) / t_seq
    print(f"[2] Filtrage + rejeu SAN->UCI, 1 cœur : {t_seq:.1f} s "
          f"({rate_seq:,.0f} parties/s)")

    # --- 3. Même chose, en parallèle ---
    workers = args.workers
    t0 = time.perf_counter()
    with mp.Pool(workers, initializer=_worker_init) as pool:
        results = pool.map(_worker_process_batch,
                           list(batched(blocks, 500)))
    t_par = time.perf_counter() - t0
    kept_par = sum(len(k) for k, _, _ in results)
    rate_par = len(blocks) / t_par
    print(f"[3] Idem sur {workers} ouvriers : {t_par:.1f} s "
          f"({rate_par:,.0f} parties/s, accélération x{rate_par/rate_seq:.1f})")

    assert kept_par == len(kept_seq), "séquentiel et parallèle divergent !"

    # --- Statistiques et extrapolation ---
    n_kept = len(kept_seq)
    retention = n_kept / len(blocks)
    lengths = [len(g.split()) for g in kept_seq]
    mean_len = sum(lengths) / len(lengths) if lengths else 0

    print(f"\n--- Sur cet échantillon de {len(blocks):,} parties ---")
    print(f"  conservées        : {n_kept:,} ({retention:.1%})")
    print(f"  longueur moyenne  : {mean_len:.1f} demi-coups")
    print(f"  tokens par partie lue : {mean_len * retention:.1f}")
    print(f"  motifs de rejet :")
    for reason, count in rejected.most_common():
        print(f"    {reason:<20} {count:>8,} ({count/len(blocks):5.1%})")

    target_tokens = args.target_tokens
    tokens_per_read_game = mean_len * retention
    if tokens_per_read_game > 0:
        games_needed = target_tokens / tokens_per_read_game
        games_kept_needed = games_needed * retention
        print(f"\n--- Extrapolation pour {target_tokens/1e6:.0f} M de tokens ---")
        print(f"  parties à lire      : {games_needed/1e6:.1f} M")
        print(f"  parties conservées  : {games_kept_needed/1e6:.1f} M")
        print(f"  durée 1 cœur        : {games_needed/rate_seq/3600:.1f} h")
        print(f"  durée {workers} ouvriers : "
              f"{games_needed/rate_par/3600:.1f} h")

    report = {
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "methode": ("lecture des N premières parties du dump, chronométrage "
                    "séparé de la décompression, du traitement séquentiel et "
                    "du traitement parallèle sur le même échantillon"),
        "echantillon_parties_lues": len(blocks),
        "cpu_workers": workers,
        "decompression_parties_s": round(len(blocks) / t_read, 1),
        "sequentiel_parties_s": round(rate_seq, 1),
        "parallele_parties_s": round(rate_par, 1),
        "acceleration": round(rate_par / rate_seq, 2),
        "taux_retention": round(retention, 4),
        "longueur_moyenne_demi_coups": round(mean_len, 2),
        "rejets": dict(rejected),
        "extrapolation_target_tokens": target_tokens,
        "extrapolation_parties_a_lire_millions": round(games_needed / 1e6, 2),
        "extrapolation_heures_sequentiel": round(games_needed / rate_seq / 3600, 2),
        "extrapolation_heures_parallele": round(games_needed / rate_par / 3600, 2),
    }
    with open(args.report, "w") as f:
        json.dump(report, f, indent=2, ensure_ascii=False)
    print(f"\nRapport écrit dans {args.report}")


# --------------------------------------------------------------------------
# Commande : parse
# --------------------------------------------------------------------------

def cmd_parse(args):
    """Parcourt tout le dump et écrit le corpus UCI."""
    t_start = time.perf_counter()
    n_read = 0
    n_kept = 0
    rejected = Counter()
    length_hist = Counter()

    os.makedirs(os.path.dirname(os.path.abspath(args.out)), exist_ok=True)

    print(f"Entrée  : {args.input}")
    print(f"Sortie  : {args.out}")
    print(f"Ouvriers: {args.workers}")
    print(f"Objectif: {args.target_tokens/1e6:.0f} M de tokens "
          f"({'illimité' if args.max_games == 0 else f'{args.max_games:,} parties max'})\n",
          flush=True)

    n_tokens = 0
    stop = False
    dump_exhausted = True

    # On traite le flux par fenêtres plutôt qu'en le passant d'un bloc au pool.
    # `Pool.imap` consomme son itérable d'entrée aussi vite qu'il le peut et
    # met en file d'attente *toutes* les tâches : sur 92 millions de parties,
    # ça représente une centaine de gigaoctets de texte en RAM et un processus
    # tué par l'OOM killer au bout de quelques minutes. En découpant en
    # fenêtres de taille fixe traitées l'une après l'autre, l'empreinte mémoire
    # reste bornée par la fenêtre — ici quelques dizaines de mégaoctets.
    window_batches = args.window // args.batch_size

    with open(args.out, "w") as out_fh, \
         mp.Pool(args.workers, initializer=_worker_init) as pool:

        batch_stream = batched(iter_game_blocks(args.input), args.batch_size)

        while not stop:
            window = []
            for batch in batch_stream:
                window.append(batch)
                if len(window) >= window_batches:
                    break
            if not window:
                break                      # flux épuisé
            if len(window) < window_batches:
                dump_exhausted = True      # dernière fenêtre, incomplète

            for kept, rej, n_blocks in pool.map(_worker_process_batch, window):
                n_read += n_blocks
                n_kept += len(kept)
                rejected.update(rej)
                for game in kept:
                    n_moves = game.count(" ") + 1
                    length_hist[n_moves] += 1
                    n_tokens += n_moves
                    out_fh.write(game)
                    out_fh.write("\n")

            elapsed = time.perf_counter() - t_start
            print(f"  {n_read/1e6:6.2f} M lues | {n_kept/1e6:5.2f} M gardées "
                  f"({n_kept/max(n_read,1):5.1%}) | {n_tokens/1e6:7.1f} M tokens "
                  f"| {n_read/elapsed:,.0f} parties/s "
                  f"| {elapsed/60:5.1f} min", flush=True)

            if args.target_tokens and n_tokens >= args.target_tokens:
                print(f"\nObjectif de {args.target_tokens/1e6:.0f} M tokens atteint.")
                dump_exhausted = False
                stop = True
            if args.max_games and n_read >= args.max_games:
                dump_exhausted = False
                stop = True

    elapsed = time.perf_counter() - t_start
    mean_len = n_tokens / n_kept if n_kept else 0

    stats = {
        "dump_epuise_avant_objectif": dump_exhausted,
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "source": os.path.basename(args.input),
        "parties_lues": n_read,
        "parties_conservees": n_kept,
        "taux_retention": round(n_kept / max(n_read, 1), 4),
        "tokens_coups_total": n_tokens,
        "longueur_moyenne_demi_coups": round(mean_len, 2),
        "duree_secondes": round(elapsed, 1),
        "debit_parties_s": round(n_read / elapsed, 1),
        "workers": args.workers,
        "rejets": dict(rejected),
        "histogramme_longueurs": dict(sorted(length_hist.items())),
        "criteres": {
            "elo_min": ELO_MIN, "elo_max": ELO_MAX,
            "min_demi_coups": MIN_PLIES, "max_demi_coups": MAX_PLIES,
            "cadences_exclues": list(EXCLUDED_SPEEDS),
            "terminaison_requise": REQUIRED_TERMINATION,
        },
    }
    with open(args.stats, "w") as f:
        json.dump(stats, f, indent=2, ensure_ascii=False)

    print(f"\n=== Terminé en {elapsed/60:.1f} min ===")
    print(f"  parties lues       : {n_read:,}")
    print(f"  parties conservées : {n_kept:,} ({n_kept/max(n_read,1):.1%})")
    print(f"  tokens de coups    : {n_tokens:,}")
    print(f"  longueur moyenne   : {mean_len:.1f} demi-coups")
    print(f"  statistiques -> {args.stats}")


# --------------------------------------------------------------------------
# Commande : encode
# --------------------------------------------------------------------------

def cmd_encode(args):
    """Encode le corpus texte en deux .bin uint16 (train et val).

    Pourquoi uint16 : le vocabulaire fait 1971 entrées, qui tiennent largement
    sur 16 bits. Stocker en uint16 plutôt qu'en int32 divise par deux la taille
    du fichier et donc le volume à faire transiter depuis le disque pendant
    l'entraînement. À 800 M de tokens, c'est 1.6 Go au lieu de 3.2 Go — assez
    petit pour que le cache disque du système garde tout en RAM.

    Chaque partie est encodée <bos> coup1 coup2 ... coupN <eos>, et toutes les
    parties sont concaténées bout à bout dans un seul tableau plat. C'est le
    format le plus simple à échantillonner : pour un batch, on tire des
    positions au hasard dans le tableau et on lit 256 tokens consécutifs.
    """
    with open(args.vocab) as f:
        vocab = json.load(f)
    stoi = vocab["stoi"]
    bos, eos = vocab["bos_id"], vocab["eos_id"]

    print(f"Vocabulaire : {vocab['vocab_size']} tokens")
    print(f"Corpus      : {args.games}")

    # Premier passage : compter les tokens pour dimensionner le memmap.
    print("Passage 1/2 — comptage...", flush=True)
    t0 = time.perf_counter()
    n_games = 0
    n_tokens = 0
    with open(args.games) as f:
        for line in f:
            n_moves = line.count(" ") + 1
            n_tokens += n_moves + 2   # <bos> et <eos>
            n_games += 1
    print(f"  {n_games:,} parties, {n_tokens:,} tokens "
          f"({time.perf_counter()-t0:.1f} s)")

    n_val_games = max(1, int(n_games * args.val_fraction))
    n_train_games = n_games - n_val_games
    print(f"  découpage : {n_train_games:,} parties d'entraînement, "
          f"{n_val_games:,} de validation ({args.val_fraction:.1%})")

    # Le découpage se fait par *partie* et non par token : une partie ne doit
    # jamais être coupée entre train et val, sinon le modèle aurait vu le début
    # d'une partie qu'on lui demande ensuite de prédire. C'est une fuite de
    # données classique, et sournoise parce qu'elle améliore la loss de
    # validation sans améliorer le modèle.
    rng = np.random.default_rng(args.seed)
    is_val = np.zeros(n_games, dtype=bool)
    is_val[rng.choice(n_games, size=n_val_games, replace=False)] = True

    print("Passage 2/2 — encodage...", flush=True)
    t0 = time.perf_counter()

    train_path = os.path.join(args.outdir, "train.bin")
    val_path = os.path.join(args.outdir, "val.bin")
    # On alloue large puis on tronque : on connaît n_tokens exactement, mais on
    # ne sait pas encore comment il se répartit entre train et val.
    train_buf = np.memmap(train_path, dtype=np.uint16, mode="w+", shape=(n_tokens,))
    val_buf = np.memmap(val_path, dtype=np.uint16, mode="w+", shape=(n_tokens,))

    # On conserve aussi les parties de validation en clair. L'évaluation de la
    # phase 4 en a besoin : pour vérifier la légalité d'un coup il faut rejouer
    # la position sur un échiquier, ce qui demande les coups en UCI et non
    # leurs identifiants numériques.
    val_txt_path = os.path.join(args.outdir, "val_games.txt")

    ti = vi = 0
    unknown = Counter()
    with open(args.games) as f, open(val_txt_path, "w") as val_txt:
        for idx, line in enumerate(f):
            moves = line.split()
            try:
                ids = [stoi[m] for m in moves]
            except KeyError as e:
                unknown[str(e)] += 1
                continue
            seq = np.fromiter([bos] + ids + [eos], dtype=np.uint16,
                              count=len(ids) + 2)
            if is_val[idx]:
                val_buf[vi:vi + len(seq)] = seq
                vi += len(seq)
                val_txt.write(line if line.endswith("\n") else line + "\n")
            else:
                train_buf[ti:ti + len(seq)] = seq
                ti += len(seq)

            if (idx + 1) % 1_000_000 == 0:
                print(f"  {idx+1:,} parties encodées "
                      f"({time.perf_counter()-t0:.0f} s)", flush=True)

    train_buf.flush()
    val_buf.flush()
    del train_buf, val_buf

    # Troncature aux tailles réelles
    with open(train_path, "r+b") as f:
        f.truncate(ti * 2)
    with open(val_path, "r+b") as f:
        f.truncate(vi * 2)

    assert not unknown, f"coups absents du vocabulaire : {unknown}"

    stats = {
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "vocab_size": vocab["vocab_size"],
        "parties_total": n_games,
        "parties_train": n_train_games,
        "parties_val": n_val_games,
        "tokens_train": int(ti),
        "tokens_val": int(vi),
        "tokens_total": int(ti + vi),
        "val_fraction_parties": args.val_fraction,
        "seed_split": args.seed,
        "octets_train": int(ti * 2),
        "octets_val": int(vi * 2),
        "dtype": "uint16",
        "format": "<bos> coups... <eos>, parties concaténées bout à bout",
        "val_games_txt": os.path.abspath(val_txt_path),
    }
    with open(os.path.join(args.outdir, "encode_stats.json"), "w") as f:
        json.dump(stats, f, indent=2, ensure_ascii=False)

    print(f"\n=== Encodage terminé en {(time.perf_counter()-t0)/60:.1f} min ===")
    print(f"  train.bin : {ti:,} tokens ({ti*2/1e9:.2f} Go)")
    print(f"  val.bin   : {vi:,} tokens ({vi*2/1e9:.2f} Go)")


# --------------------------------------------------------------------------

def main():
    p = argparse.ArgumentParser(description=__doc__,
                                formatter_class=argparse.RawDescriptionHelpFormatter)
    sub = p.add_subparsers(dest="cmd", required=True)

    default_input = "data/lichess_db_standard_rated_2026-07.pgn.zst"

    b = sub.add_parser("benchmark", help="mesure le débit de parsing")
    b.add_argument("--input", default=default_input)
    b.add_argument("--games", type=int, default=50_000)
    b.add_argument("--workers", type=int, default=max(1, os.cpu_count() - 2))
    b.add_argument("--target-tokens", type=float, default=800e6)
    b.add_argument("--report", default="logs/phase1_benchmark.json")
    b.set_defaults(func=cmd_benchmark)

    q = sub.add_parser("parse", help="traite tout le dump")
    q.add_argument("--input", default=default_input)
    q.add_argument("--out", default="data/games_uci.txt")
    q.add_argument("--stats", default="logs/phase1_parse_stats.json")
    q.add_argument("--workers", type=int, default=max(1, os.cpu_count() - 2))
    q.add_argument("--batch-size", type=int, default=500)
    q.add_argument("--window", type=int, default=200_000,
                   help="parties chargées en RAM à la fois (borne mémoire)")
    q.add_argument("--target-tokens", type=float, default=800e6)
    q.add_argument("--max-games", type=int, default=0)
    q.set_defaults(func=cmd_parse)

    e = sub.add_parser("encode", help="encode en .bin uint16")
    e.add_argument("--games", default="data/games_uci.txt")
    e.add_argument("--vocab", default="data/vocab.json")
    e.add_argument("--outdir", default="data")
    e.add_argument("--val-fraction", type=float, default=0.01)
    e.add_argument("--seed", type=int, default=1337)
    e.set_defaults(func=cmd_encode)

    args = p.parse_args()
    args.target_tokens = int(getattr(args, "target_tokens", 0) or 0)
    args.func(args)


if __name__ == "__main__":
    main()
model.py392 lines

The Transformer: pre-norm, RMSNorm, RoPE, SwiGLU. Every architectural choice is commented in the file.

"""Le modèle : un Transformer décodeur, écrit en PyTorch pur.

L'idée centrale du projet tient en une phrase : on ne code aucune règle du jeu.
Le modèle reçoit une suite de coups et doit prédire le suivant, exactement comme
un modèle de langage prédit le mot suivant. Il n'a pas d'échiquier en mémoire,
pas de notion de pièce, pas de vérificateur de légalité. S'il finit par jouer
des coups légaux, c'est uniquement parce que c'est le seul moyen de bien
prédire des parties jouées par des humains.

L'architecture est celle des modèles récents (Llama et ses descendants) plutôt
que celle du GPT-2 d'origine. Quatre choix, expliqués un par un plus bas :
pre-norm, RMSNorm, RoPE et SwiGLU.
"""

import math
from dataclasses import dataclass

import torch
import torch.nn as nn
import torch.nn.functional as F


@dataclass
class ModelConfig:
    vocab_size: int = 1971      # 1968 coups UCI + <pad> <bos> <eos>
    block_size: int = 256       # contexte, en coups
    n_layer: int = 16
    n_head: int = 8
    n_embd: int = 512
    mlp_hidden: int = 1408      # multiple de 64, ~8/3 x n_embd (voir SwiGLU)
    rope_theta: float = 10000.0
    dropout: float = 0.0        # on a bien plus de données que de paramètres
    tie_embeddings: bool = True


# ---------------------------------------------------------------------------
# RMSNorm
# ---------------------------------------------------------------------------

class RMSNorm(nn.Module):
    """Normalisation par la racine de la moyenne des carrés.

    La LayerNorm classique fait deux choses : elle recentre les valeurs autour
    de zéro, puis les remet à une échelle unitaire. La RMSNorm ne fait que la
    seconde. Il s'avère que le recentrage n'apporte presque rien en pratique,
    alors que le calculer coûte une passe supplémentaire sur les données.
    On garde donc uniquement la mise à l'échelle.

    L'intuition : ce qui compte pour la stabilité de l'entraînement, c'est que
    les valeurs qui circulent dans le réseau gardent une taille comparable
    d'une couche à l'autre. Savoir si elles sont centrées sur zéro ou sur trois
    est secondaire.
    """

    def __init__(self, dim: int, eps: float = 1e-6):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))

    def forward(self, x):
        # Le calcul se fait en float32 même quand x est en bf16 : une somme de
        # carrés sur 512 valeurs perd trop de précision en 16 bits.
        dtype = x.dtype
        x = x.float()
        x = x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
        return (x.to(dtype)) * self.weight


# ---------------------------------------------------------------------------
# RoPE — Rotary Position Embedding
# ---------------------------------------------------------------------------

def build_rope_cache(block_size: int, head_dim: int, theta: float, device, dtype):
    """Pré-calcule les cosinus et sinus de rotation pour chaque position.

    Comment un Transformer sait-il dans quel ordre sont les coups ? Le mécanisme
    d'attention, seul, n'en a aucune idée : il traite ses entrées comme un sac
    d'éléments. Il faut donc lui injecter la position.

    L'approche historique ajoutait un vecteur de position à chaque token. RoPE
    fait autrement : au lieu d'ajouter une information de position, il **fait
    tourner** les vecteurs requête et clé d'un angle proportionnel à leur
    position. Les dimensions sont prises deux par deux et chaque paire subit une
    rotation dans son plan.

    L'élégance de la chose : quand on calcule ensuite le produit scalaire entre
    une requête en position m et une clé en position n, le résultat ne dépend
    plus que de la différence m - n. L'attention devient donc naturellement
    sensible aux distances relatives — « le coup d'il y a trois demi-coups »
    plutôt que « le coup numéro 47 ». Pour des parties d'échecs c'est
    exactement ce qu'on veut : un motif tactique a la même signification qu'il
    survienne au coup 10 ou au coup 40.
    """
    # Chaque paire de dimensions tourne à une vitesse différente : les
    # premières tournent vite (elles encodent les positions proches), les
    # dernières tournent lentement (elles encodent la position globale).
    inv_freq = 1.0 / (theta ** (torch.arange(0, head_dim, 2, device=device).float()
                                / head_dim))
    positions = torch.arange(block_size, device=device).float()
    angles = torch.outer(positions, inv_freq)          # (T, head_dim/2)
    return angles.cos().to(dtype), angles.sin().to(dtype)


def apply_rope(x, cos, sin):
    """Applique la rotation à un tenseur (B, n_head, T, head_dim)."""
    # On sépare les dimensions en deux moitiés, qui jouent le rôle des parties
    # réelle et imaginaire d'un nombre complexe.
    x1, x2 = x.chunk(2, dim=-1)
    cos = cos[None, None, :, :]
    sin = sin[None, None, :, :]
    return torch.cat([x1 * cos - x2 * sin,
                      x2 * cos + x1 * sin], dim=-1)


# ---------------------------------------------------------------------------
# Attention
# ---------------------------------------------------------------------------

class CausalSelfAttention(nn.Module):
    """Attention multi-tête causale.

    « Causale » signifie que le coup en position t ne peut regarder que les
    coups 0 à t. C'est ce qui rend la tâche non triviale : sans ce masque, le
    modèle verrait le coup qu'on lui demande de prédire.
    """

    def __init__(self, cfg: ModelConfig):
        super().__init__()
        assert cfg.n_embd % cfg.n_head == 0
        self.n_head = cfg.n_head
        self.head_dim = cfg.n_embd // cfg.n_head
        # Les trois projections q, k, v sont fusionnées en une seule matrice :
        # un gros produit matriciel est plus efficace que trois petits.
        self.qkv = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=False)
        self.proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=False)
        self.dropout = cfg.dropout

    def forward(self, x, cos, sin):
        B, T, C = x.shape
        q, k, v = self.qkv(x).split(C, dim=2)
        # (B, T, C) -> (B, n_head, T, head_dim)
        q = q.view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        k = k.view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        v = v.view(B, T, self.n_head, self.head_dim).transpose(1, 2)

        q = apply_rope(q, cos[:T], sin[:T])
        k = apply_rope(k, cos[:T], sin[:T])

        # PyTorch choisit tout seul l'implémentation la plus rapide disponible
        # (FlashAttention sur ces cartes), qui ne matérialise jamais la matrice
        # d'attention T x T en mémoire.
        y = F.scaled_dot_product_attention(
            q, k, v,
            dropout_p=self.dropout if self.training else 0.0,
            is_causal=True,
        )
        y = y.transpose(1, 2).contiguous().view(B, T, C)
        return self.proj(y)


# ---------------------------------------------------------------------------
# SwiGLU
# ---------------------------------------------------------------------------

class SwiGLU(nn.Module):
    """Le bloc de calcul « par position » du Transformer.

    Un MLP classique fait : projection vers le haut, non-linéarité, projection
    vers le bas. SwiGLU ajoute un mécanisme de porte : on calcule deux
    projections vers le haut, on passe la première dans une fonction SiLU, et
    on **multiplie** les deux résultats terme à terme. La seconde projection
    agit donc comme un robinet qui laisse passer plus ou moins chaque
    dimension de la première.

    L'intuition : plutôt que d'appliquer la même transformation à tout ce qui
    passe, le réseau apprend à décider, dimension par dimension et en fonction
    de l'entrée, ce qui mérite d'être transmis à la couche suivante.

    Ça coûte une troisième matrice là où un MLP classique en a deux. C'est
    pourquoi on réduit la dimension cachée d'un facteur 2/3 par rapport aux 4x
    habituels : 4 x 2/3 = 8/3, soit ici 512 x 8/3 = 1365, arrondi à 1408 pour
    tomber sur un multiple de 64 (les cœurs tensoriels des GPU aiment ça).
    Le nombre de paramètres reste ainsi comparable à celui d'un MLP classique.
    """

    def __init__(self, cfg: ModelConfig):
        super().__init__()
        self.gate = nn.Linear(cfg.n_embd, cfg.mlp_hidden, bias=False)
        self.up = nn.Linear(cfg.n_embd, cfg.mlp_hidden, bias=False)
        self.down = nn.Linear(cfg.mlp_hidden, cfg.n_embd, bias=False)

    def forward(self, x):
        return self.down(F.silu(self.gate(x)) * self.up(x))


# ---------------------------------------------------------------------------
# Bloc
# ---------------------------------------------------------------------------

class Block(nn.Module):
    """Un étage du Transformer, en pre-norm.

    « Pre-norm » veut dire qu'on normalise **avant** chaque sous-couche, et non
    après. La différence paraît cosmétique, elle ne l'est pas : en pre-norm, le
    chemin qui traverse les connexions résiduelles ne rencontre aucune
    normalisation. Le gradient peut donc remonter de la sortie jusqu'à la
    première couche sans être altéré, ce qui rend les réseaux profonds
    entraînables sans phase de chauffe acrobatique. C'est ce qui a permis
    d'empiler des dizaines de couches de façon fiable.
    """

    def __init__(self, cfg: ModelConfig):
        super().__init__()
        self.norm_attn = RMSNorm(cfg.n_embd)
        self.attn = CausalSelfAttention(cfg)
        self.norm_mlp = RMSNorm(cfg.n_embd)
        self.mlp = SwiGLU(cfg)

    def forward(self, x, cos, sin):
        x = x + self.attn(self.norm_attn(x), cos, sin)
        x = x + self.mlp(self.norm_mlp(x))
        return x


# ---------------------------------------------------------------------------
# Le modèle
# ---------------------------------------------------------------------------

class ChessGPT(nn.Module):

    def __init__(self, cfg: ModelConfig):
        super().__init__()
        self.cfg = cfg
        self.tok_emb = nn.Embedding(cfg.vocab_size, cfg.n_embd)
        self.blocks = nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)])
        self.norm_final = RMSNorm(cfg.n_embd)
        self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)

        if cfg.tie_embeddings:
            # La matrice d'embedding et celle de sortie sont la même. C'est
            # défendable : les deux encodent « quel coup ressemble à quel
            # coup ». Ça économise un million de paramètres et ça régularise.
            self.lm_head.weight = self.tok_emb.weight

        self.apply(self._init_weights)
        # Les projections de sortie de chaque bloc sont initialisées plus
        # petit, en 1/sqrt(2 * n_layer). Sans ça, la variance du flux résiduel
        # croît avec la profondeur : chaque bloc ajoute sa contribution, et
        # après seize couches le signal a explosé.
        for name, p in self.named_parameters():
            if name.endswith("proj.weight") or name.endswith("down.weight"):
                nn.init.normal_(p, mean=0.0,
                                std=0.02 / math.sqrt(2 * cfg.n_layer))

        cos, sin = build_rope_cache(cfg.block_size,
                                    cfg.n_embd // cfg.n_head,
                                    cfg.rope_theta,
                                    device="cpu", dtype=torch.float32)
        # register_buffer : ces tenseurs suivent le modèle sur le GPU et dans
        # les checkpoints, mais ne sont pas des paramètres à entraîner.
        self.register_buffer("rope_cos", cos, persistent=False)
        self.register_buffer("rope_sin", sin, persistent=False)

    @staticmethod
    def _init_weights(module):
        if isinstance(module, nn.Linear):
            nn.init.normal_(module.weight, mean=0.0, std=0.02)
            if module.bias is not None:
                nn.init.zeros_(module.bias)
        elif isinstance(module, nn.Embedding):
            nn.init.normal_(module.weight, mean=0.0, std=0.02)

    def forward(self, idx, targets=None):
        B, T = idx.shape
        assert T <= self.cfg.block_size, \
            f"séquence de {T} tokens, contexte limité à {self.cfg.block_size}"

        x = self.tok_emb(idx)
        cos = self.rope_cos[:T]
        sin = self.rope_sin[:T]
        for block in self.blocks:
            x = block(x, cos, sin)
        x = self.norm_final(x)
        logits = self.lm_head(x)

        loss = None
        if targets is not None:
            loss = F.cross_entropy(
                logits.view(-1, logits.size(-1)),
                targets.reshape(-1),
                ignore_index=-1,
            )
        return logits, loss

    # -- comptages, utiles pour le MFU et pour la documentation --------------

    def num_params(self, non_embedding: bool = True) -> int:
        n = sum(p.numel() for p in self.parameters())
        if non_embedding:
            n -= self.tok_emb.weight.numel()
            # Si les poids sont liés, lm_head partage la même matrice et a
            # déjà été retiré ; sinon il faut le retirer aussi.
            if not self.cfg.tie_embeddings:
                n -= self.lm_head.weight.numel()
        return n

    def flops_per_token(self) -> float:
        """Estime le coût en FLOP d'un token, passe avant + arrière.

        Méthode dite « de Chinchilla » : on compte 6 FLOP par paramètre et par
        token (2 pour la passe avant, 4 pour la passe arrière), auxquels on
        ajoute le coût de l'attention, qui lui dépend de la longueur du
        contexte et n'est donc pas capté par le comptage de paramètres.

        Ce chiffre sert à calculer le MFU (Model FLOPs Utilization) : la part
        de la puissance brute du GPU qu'on exploite réellement. Un MFU de 40 %
        sur ce genre de modèle est un bon score ; en dessous de 20 %, il y a
        un goulot à chercher.
        """
        cfg = self.cfg
        n = self.num_params(non_embedding=True)
        # 6 * N pour les couches denses
        flops = 6 * n
        # Attention : les deux produits QK^T et (attn @ V), passe avant et
        # arrière, soit 12 FLOP par (couche, tête, position vue).
        flops += 12 * cfg.n_layer * cfg.n_head * (cfg.n_embd // cfg.n_head) * cfg.block_size
        return flops

    @torch.no_grad()
    def generate(self, idx, max_new_tokens, temperature=1.0, top_k=None,
                 allowed_mask_fn=None):
        """Génère des coups un par un.

        `allowed_mask_fn` permet, si on le souhaite, de restreindre les coups
        possibles à chaque étape (c'est ce que fera engine.py pour garantir la
        légalité). Laissé à None, le modèle génère librement — c'est le mode
        utilisé pour mesurer le taux de coups légaux, qui est le chiffre
        central de ce projet.
        """
        for _ in range(max_new_tokens):
            idx_cond = idx[:, -self.cfg.block_size:]
            logits, _ = self(idx_cond)
            logits = logits[:, -1, :]

            if allowed_mask_fn is not None:
                mask = allowed_mask_fn(idx)
                logits = logits.masked_fill(~mask, float("-inf"))

            if temperature == 0.0:
                next_id = logits.argmax(dim=-1, keepdim=True)
            else:
                logits = logits / temperature
                if top_k is not None:
                    v, _ = torch.topk(logits, min(top_k, logits.size(-1)))
                    logits = logits.masked_fill(logits < v[:, [-1]],
                                                float("-inf"))
                probs = F.softmax(logits, dim=-1)
                next_id = torch.multinomial(probs, num_samples=1)

            idx = torch.cat((idx, next_id), dim=1)
        return idx


if __name__ == "__main__":
    cfg = ModelConfig()
    model = ChessGPT(cfg)
    total = sum(p.numel() for p in model.parameters())
    print(f"Configuration : {cfg.n_layer} couches, d={cfg.n_embd}, "
          f"{cfg.n_head} têtes, MLP caché {cfg.mlp_hidden}, "
          f"contexte {cfg.block_size}")
    print(f"  paramètres totaux        : {total:,}")
    print(f"  paramètres non-embedding : {model.num_params():,}")
    print(f"  FLOP par token (fwd+bwd) : {model.flops_per_token():,.0f}")

    # Test de passe avant avec la vraie tâche : prédire le token *suivant*.
    # Attention au piège : évaluer le modèle sur `targets = inputs` (prédire le
    # token courant) donne une loss bien inférieure au hasard dès
    # l'initialisation. Ce n'est pas un bug du modèle mais une propriété des
    # embeddings liés : la connexion résiduelle transporte l'embedding du token
    # jusqu'à la sortie, où lm_head — qui *est* la matrice d'embedding — le
    # reconnaît. Le réseau sait donc recopier son entrée avant tout
    # entraînement. Il ne sait rien prédire pour autant.
    seq = torch.randint(0, cfg.vocab_size, (2, cfg.block_size + 1))
    inputs, targets = seq[:, :-1], seq[:, 1:]
    logits, loss = model(inputs, targets)
    print(f"  test de passe avant      : logits {tuple(logits.shape)}, "
          f"loss {loss.item():.4f}")
    print(f"  loss attendue au hasard  : {math.log(cfg.vocab_size):.4f}")

    _, loss_identity = model(inputs, inputs)
    print(f"  (contrôle) loss si la cible est l'entrée elle-même : "
          f"{loss_identity.item():.4f}")
train.py572 lines

The training loop, the checkpoints, resumption, and the overfitting mode that acts as a safeguard.

"""Boucle d'entraînement, en PyTorch pur.

Aucun framework d'entraînement : ni HuggingFace Trainer, ni Lightning. Chaque
ligne est lisible, ce qui est précisément l'intérêt du projet — on veut pouvoir
suivre ce qui se passe plutôt que de faire confiance à une abstraction.

Trois modes :

    # Test de surapprentissage : 100 parties, la loss doit tomber vers zéro.
    # C'est le garde-fou obligatoire avant le run long.
    python train.py --overfit

    # Entraînement complet
    python train.py --run-name run1

    # Reprise après interruption
    python train.py --run-name run1 --resume
"""

import argparse
import csv
import json
import math
import os
import time
from contextlib import nullcontext

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import numpy as np
import torch

from model import ChessGPT, ModelConfig

# TFLOPS bf16 réellement mesurés sur la carte au 2026-08-04 (plafond 280 W).
# On calcule le MFU par rapport à ce chiffre et non par rapport à la fiche
# technique du constructeur : c'est la seule référence honnête, puisque c'est
# la performance dont on dispose réellement.
MEASURED_PEAK_TFLOPS_BF16 = 62.86


# ---------------------------------------------------------------------------
# Données
# ---------------------------------------------------------------------------

class MemmapDataset:
    """Échantillonne des fenêtres de `block_size` tokens dans un .bin uint16.

    Le corpus est un unique tableau plat où toutes les parties sont
    concaténées, séparées par <eos><bos>. Pour former un batch, on tire des
    positions au hasard et on lit les tokens consécutifs à partir de là.

    Une fenêtre peut donc chevaucher deux parties. C'est voulu : avec une
    moyenne de 73.5 tokens par partie et un contexte de 256, découper
    proprement une partie par séquence gaspillerait 70 % du calcul en
    remplissage. Le modèle apprend simplement que <eos><bos> signifie « tout
    ce qui précède ne compte plus » — ce qu'il doit de toute façon apprendre
    pour savoir reconnaître une fin de partie.
    """

    def __init__(self, path: str, block_size: int, device: str):
        self.path = path
        self.block_size = block_size
        self.device = device
        # On rouvre le memmap à chaque époque plutôt que de garder une
        # référence : c'est la recommandation de numpy pour éviter une fuite
        # de descripteurs sur les fichiers volumineux.
        self.n_tokens = os.path.getsize(path) // 2

    def get_batch(self, batch_size: int, generator=None):
        data = np.memmap(self.path, dtype=np.uint16, mode="r")
        # -1 car il faut un token de plus pour la cible décalée
        high = len(data) - self.block_size - 1
        ix = torch.randint(high, (batch_size,), generator=generator)
        x = torch.stack([
            torch.from_numpy(data[i:i + self.block_size].astype(np.int64))
            for i in ix])
        y = torch.stack([
            torch.from_numpy(data[i + 1:i + 1 + self.block_size].astype(np.int64))
            for i in ix])
        # pin_memory + non_blocking : le transfert vers le GPU se recouvre avec
        # le calcul du batch précédent au lieu de le bloquer.
        return (x.pin_memory().to(self.device, non_blocking=True),
                y.pin_memory().to(self.device, non_blocking=True))


class OverfitDataset:
    """Un jeu minuscule gardé entièrement en mémoire GPU.

    Sert au test de surapprentissage : si le modèle n'arrive pas à apprendre
    par cœur cent parties, c'est qu'il y a un bug quelque part — dans le
    modèle, dans les données, ou dans la boucle. Autant le découvrir en trois
    minutes plutôt qu'après une nuit de calcul.
    """

    def __init__(self, games_path: str, vocab_path: str, block_size: int,
                 n_games: int, device: str):
        with open(vocab_path) as f:
            vocab = json.load(f)
        stoi, bos, eos = vocab["stoi"], vocab["bos_id"], vocab["eos_id"]

        ids = []
        with open(games_path) as f:
            for i, line in enumerate(f):
                if i >= n_games:
                    break
                ids.append(bos)
                ids.extend(stoi[m] for m in line.split())
                ids.append(eos)

        self.data = torch.tensor(ids, dtype=torch.int64, device=device)
        self.block_size = block_size
        self.device = device
        self.n_tokens = len(ids)
        self.n_games = min(n_games, i + 1)

    def get_batch(self, batch_size: int, generator=None):
        high = len(self.data) - self.block_size - 1
        ix = torch.randint(high, (batch_size,), device=self.device)
        x = torch.stack([self.data[i:i + self.block_size] for i in ix])
        y = torch.stack([self.data[i + 1:i + 1 + self.block_size] for i in ix])
        return x, y


# ---------------------------------------------------------------------------
# Planification du taux d'apprentissage
# ---------------------------------------------------------------------------

def get_lr(step: int, cfg) -> float:
    """Chauffe linéaire puis décroissance en cosinus.

    Pourquoi une chauffe : au tout début, les gradients sont énormes et
    désordonnés. Appliquer immédiatement le taux d'apprentissage nominal fait
    diverger le modèle, ou pire, le fait converger vers une solution médiocre
    dont il ne ressortira pas. On monte donc progressivement.

    Pourquoi un cosinus ensuite : en fin d'entraînement on veut de tout petits
    pas, pour se poser dans un minimum plutôt que de rebondir autour. Le
    cosinus donne une décroissance douce au début et de plus en plus marquée,
    ce qui empiriquement bat la décroissance linéaire.
    """
    if step < cfg.warmup_steps:
        return cfg.lr * (step + 1) / cfg.warmup_steps
    if step >= cfg.max_steps:
        return cfg.lr_min
    progress = (step - cfg.warmup_steps) / max(1, cfg.max_steps - cfg.warmup_steps)
    coeff = 0.5 * (1.0 + math.cos(math.pi * progress))
    return cfg.lr_min + coeff * (cfg.lr - cfg.lr_min)


# ---------------------------------------------------------------------------
# Optimiseur
# ---------------------------------------------------------------------------

def build_optimizer(model, lr, weight_decay, betas, device_type):
    """AdamW, avec weight decay uniquement sur les matrices.

    Le weight decay tire les poids vers zéro pour limiter le surapprentissage.
    On l'applique aux matrices de projection, mais pas aux vecteurs de gain des
    normalisations ni aux biais : ces paramètres-là contrôlent une échelle, et
    les pousser vers zéro revient à éteindre progressivement des couches
    entières. C'est une distinction que beaucoup d'implémentations oublient.
    """
    decay, no_decay = [], []
    for name, p in model.named_parameters():
        if not p.requires_grad:
            continue
        (decay if p.dim() >= 2 else no_decay).append(p)

    groups = [
        {"params": decay, "weight_decay": weight_decay},
        {"params": no_decay, "weight_decay": 0.0},
    ]
    # fused=True fait tourner la mise à jour d'Adam dans un seul noyau CUDA
    # au lieu d'un par tenseur : sur un modèle à 300 tenseurs, ça compte.
    use_fused = device_type == "cuda"
    opt = torch.optim.AdamW(groups, lr=lr, betas=betas, fused=use_fused)
    return opt, len(decay), len(no_decay)


# ---------------------------------------------------------------------------
# Évaluation
# ---------------------------------------------------------------------------

@torch.no_grad()
def estimate_loss(model, dataset, batch_size, n_batches, ctx):
    model.eval()
    losses = torch.zeros(n_batches)
    for i in range(n_batches):
        x, y = dataset.get_batch(batch_size)
        with ctx:
            _, loss = model(x, y)
        losses[i] = loss.item()
    model.train()
    return losses.mean().item()


# ---------------------------------------------------------------------------
# Checkpoints
# ---------------------------------------------------------------------------

def save_snapshot(path, raw_model, step, tokens_seen, best_val):
    """Instantané léger, pour l'évaluation des checkpoints intermédiaires.

    Différence avec un checkpoint de reprise : on ne sauvegarde **que** les
    poids, sans l'état de l'optimiseur. Celui-ci pèse deux fois le modèle (les
    deux moments d'Adam) et ne sert qu'à reprendre un entraînement, jamais à
    évaluer. Un instantané fait donc 210 Mo au lieu de 630.

    Ces instantanés sont numérotés et ne s'écrasent pas, contrairement au
    checkpoint de reprise. C'est ce qui permet de tracer a posteriori la
    courbe « taux de coups légaux en fonction des tokens vus » — le graphique
    le plus intéressant du projet, et qu'on ne peut pas reconstituer si on a
    écrasé les états intermédiaires.
    """
    tmp = path + ".tmp"
    torch.save({
        "model": raw_model.state_dict(),
        "step": step,
        "tokens_seen": tokens_seen,
        "model_config": raw_model.cfg.__dict__,
        "best_val": best_val,
        "snapshot": True,
    }, tmp)
    os.replace(tmp, path)


def save_checkpoint(path, raw_model, optimizer, step, tokens_seen, cfg,
                    best_val, elapsed):
    """Sauvegarde tout ce qu'il faut pour reprendre exactement où on en est.

    Sauvegarder les poids ne suffit pas : sans l'état de l'optimiseur (les deux
    moments d'Adam, qui sont une moyenne mobile des gradients passés), une
    reprise repart avec un optimiseur amnésique et la loss fait un bond visible
    pendant quelques centaines de pas.

    On écrit d'abord dans un fichier temporaire, puis on renomme. Le renommage
    est atomique sur un système de fichiers POSIX : si le processus est tué
    pendant l'écriture, on garde le checkpoint précédent intact plutôt que de
    se retrouver avec un fichier tronqué.
    """
    tmp = path + ".tmp"
    torch.save({
        "model": raw_model.state_dict(),
        "optimizer": optimizer.state_dict(),
        "step": step,
        "tokens_seen": tokens_seen,
        "model_config": raw_model.cfg.__dict__,
        "train_config": vars(cfg),
        "best_val": best_val,
        "elapsed": elapsed,
        "torch_version": torch.__version__,
    }, tmp)
    os.replace(tmp, path)


# ---------------------------------------------------------------------------

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--run-name", default="run1")
    p.add_argument("--data-dir", default="data")
    p.add_argument("--out-dir", default="checkpoints")
    p.add_argument("--log-dir", default="logs")
    p.add_argument("--device", default="cuda:1",
                   help="cuda:1 est le RTX 3090 avec CUDA_DEVICE_ORDER=PCI_BUS_ID")

    # Modèle
    p.add_argument("--n-layer", type=int, default=16)
    p.add_argument("--n-head", type=int, default=8)
    p.add_argument("--n-embd", type=int, default=512)
    p.add_argument("--mlp-hidden", type=int, default=1408)
    p.add_argument("--block-size", type=int, default=256)

    # Optimisation
    p.add_argument("--batch-size", type=int, default=192,
                   help="séquences par micro-batch")
    p.add_argument("--grad-accum", type=int, default=1)
    p.add_argument("--lr", type=float, default=6e-4)
    p.add_argument("--lr-min", type=float, default=6e-5)
    p.add_argument("--warmup-steps", type=int, default=500)
    p.add_argument("--max-steps", type=int, default=16000)
    p.add_argument("--weight-decay", type=float, default=0.1)
    p.add_argument("--beta1", type=float, default=0.9)
    p.add_argument("--beta2", type=float, default=0.95)
    p.add_argument("--grad-clip", type=float, default=1.0)

    # Journalisation et checkpoints
    p.add_argument("--eval-interval", type=int, default=250)
    p.add_argument("--eval-batches", type=int, default=50)
    p.add_argument("--log-interval", type=int, default=10)
    p.add_argument("--ckpt-minutes", type=float, default=30.0)
    p.add_argument("--snapshot-every", type=int, default=1000,
                   help="instantané numéroté tous les N steps (0 pour désactiver)")
    p.add_argument("--resume", action="store_true")
    p.add_argument("--compile", dest="compile", action="store_true", default=True)
    p.add_argument("--no-compile", dest="compile", action="store_false")

    # Test de surapprentissage
    p.add_argument("--overfit", action="store_true")
    p.add_argument("--overfit-games", type=int, default=100)
    p.add_argument("--overfit-steps", type=int, default=1200)
    p.add_argument("--overfit-target", type=float, default=0.05,
                   help="loss en dessous de laquelle le test est considéré réussi")

    cfg = p.parse_args()

    torch.manual_seed(1337)
    torch.backends.cuda.matmul.allow_tf32 = True
    torch.backends.cudnn.allow_tf32 = True

    device = cfg.device
    device_type = "cuda" if device.startswith("cuda") else "cpu"
    # bf16 : même plage d'exposants que le float32, donc pas de risque de
    # débordement, et pas besoin du GradScaler qu'imposait le fp16.
    ctx = (torch.autocast(device_type=device_type, dtype=torch.bfloat16)
           if device_type == "cuda" else nullcontext())

    # Le vocabulaire est cherché d'abord dans le répertoire de données, puis
    # dans `data/`. Sans ce repli, entraîner sur un corpus rangé dans un
    # sous-répertoire (data/4mois/) échoue au démarrage parce que vocab.json
    # est resté à la racine — ce qui est exactement arrivé, après trois heures
    # de préparation de données réussies.
    for candidat in (os.path.join(cfg.data_dir, "vocab.json"),
                     os.path.join("data", "vocab.json")):
        if os.path.exists(candidat):
            with open(candidat) as f:
                vocab_size = json.load(f)["vocab_size"]
            print(f"[vocab] {candidat} ({vocab_size} tokens)")
            break
    else:
        raise SystemExit(
            f"vocab.json introuvable, ni dans {cfg.data_dir} ni dans data/")

    model_cfg = ModelConfig(
        vocab_size=vocab_size,
        block_size=cfg.block_size,
        n_layer=cfg.n_layer,
        n_head=cfg.n_head,
        n_embd=cfg.n_embd,
        mlp_hidden=cfg.mlp_hidden,
    )

    if cfg.overfit:
        cfg.max_steps = cfg.overfit_steps
        cfg.warmup_steps = min(cfg.warmup_steps, cfg.overfit_steps // 10)
        cfg.eval_interval = max(10, cfg.overfit_steps // 40)

    # --- Données ---
    if cfg.overfit:
        train_ds = OverfitDataset(
            os.path.join(cfg.data_dir, "games_uci.txt"),
            os.path.join(cfg.data_dir, "vocab.json"),
            cfg.block_size, cfg.overfit_games, device)
        val_ds = train_ds
        print(f"[données] mode surapprentissage : {train_ds.n_games} parties, "
              f"{train_ds.n_tokens:,} tokens, gardés en mémoire GPU")
    else:
        train_ds = MemmapDataset(os.path.join(cfg.data_dir, "train.bin"),
                                 cfg.block_size, device)
        val_ds = MemmapDataset(os.path.join(cfg.data_dir, "val.bin"),
                               cfg.block_size, device)
        print(f"[données] train {train_ds.n_tokens:,} tokens | "
              f"val {val_ds.n_tokens:,} tokens")

    # --- Modèle ---
    model = ChessGPT(model_cfg).to(device)
    raw_model = model          # référence non compilée, pour les checkpoints
    n_params = model.num_params()
    flops_per_token = model.flops_per_token()
    print(f"[modèle] {n_params:,} paramètres non-embedding | "
          f"{flops_per_token:,.0f} FLOP/token")

    optimizer, n_decay, n_nodecay = build_optimizer(
        model, cfg.lr, cfg.weight_decay, (cfg.beta1, cfg.beta2), device_type)
    print(f"[optim] AdamW fused | {n_decay} tenseurs avec weight decay, "
          f"{n_nodecay} sans")

    # --- Reprise ---
    step = 0
    tokens_seen = 0
    best_val = float("inf")
    elapsed_before = 0.0
    ckpt_path = os.path.join(cfg.out_dir, f"{cfg.run_name}_last.pt")
    best_path = os.path.join(cfg.out_dir, f"{cfg.run_name}_best.pt")

    if cfg.resume:
        if not os.path.exists(ckpt_path):
            raise SystemExit(f"--resume demandé mais {ckpt_path} est absent")
        print(f"[reprise] chargement de {ckpt_path}")
        ck = torch.load(ckpt_path, map_location=device, weights_only=False)
        raw_model.load_state_dict(ck["model"])
        optimizer.load_state_dict(ck["optimizer"])
        step = ck["step"]
        tokens_seen = ck["tokens_seen"]
        best_val = ck["best_val"]
        elapsed_before = ck["elapsed"]
        print(f"[reprise] step {step}, {tokens_seen:,} tokens déjà vus, "
              f"{elapsed_before/3600:.2f} h de calcul cumulées")

    if cfg.compile:
        print("[compile] torch.compile en cours (première itération lente)...",
              flush=True)
        model = torch.compile(model)

    os.makedirs(cfg.out_dir, exist_ok=True)
    os.makedirs(cfg.log_dir, exist_ok=True)

    # --- Journal CSV ---
    csv_path = os.path.join(cfg.log_dir, f"{cfg.run_name}_metrics.csv")
    new_csv = not (cfg.resume and os.path.exists(csv_path))
    csv_file = open(csv_path, "a", newline="")
    writer = csv.writer(csv_file)
    if new_csv:
        writer.writerow(["step", "tokens_vus", "loss_train", "loss_val",
                         "lr", "tokens_par_s", "mfu", "secondes_ecoulees",
                         "grad_norm"])
        csv_file.flush()

    tokens_per_step = cfg.batch_size * cfg.block_size * cfg.grad_accum
    print(f"[plan] {tokens_per_step:,} tokens/step x {cfg.max_steps:,} steps "
          f"= {tokens_per_step * cfg.max_steps / 1e6:.1f} M tokens")

    # --- Boucle ---
    t_start = time.perf_counter()
    t_last_ckpt = t_start
    t_log = time.perf_counter()
    tokens_since_log = 0
    loss_train_running = float("nan")
    overfit_reached = False

    model.train()
    while step < cfg.max_steps:
        lr = get_lr(step, cfg)
        for group in optimizer.param_groups:
            group["lr"] = lr

        optimizer.zero_grad(set_to_none=True)
        loss_accum = 0.0
        for micro in range(cfg.grad_accum):
            x, y = train_ds.get_batch(cfg.batch_size)
            with ctx:
                _, loss = model(x, y)
                loss = loss / cfg.grad_accum
            loss.backward()
            loss_accum += loss.item()

        grad_norm = torch.nn.utils.clip_grad_norm_(model.parameters(),
                                                   cfg.grad_clip)
        optimizer.step()

        step += 1
        tokens_seen += tokens_per_step
        tokens_since_log += tokens_per_step
        loss_train_running = loss_accum

        # --- Journalisation périodique ---
        if step % cfg.log_interval == 0:
            torch.cuda.synchronize(device) if device_type == "cuda" else None
            now = time.perf_counter()
            dt = now - t_log
            tok_per_s = tokens_since_log / dt
            # MFU : la part de la puissance réellement mesurée du GPU qu'on
            # exploite. Numérateur = FLOP effectivement utiles par seconde.
            mfu = (tok_per_s * flops_per_token) / (MEASURED_PEAK_TFLOPS_BF16 * 1e12)
            elapsed = elapsed_before + (now - t_start)
            print(f"step {step:>6} | loss {loss_accum:6.4f} | lr {lr:.2e} | "
                  f"{tok_per_s:>9,.0f} tok/s | MFU {mfu:5.1%} | "
                  f"{tokens_seen/1e6:7.1f} M tokens | {elapsed/60:6.1f} min",
                  flush=True)
            writer.writerow([step, tokens_seen, f"{loss_accum:.6f}", "",
                             f"{lr:.6e}", f"{tok_per_s:.1f}", f"{mfu:.4f}",
                             f"{elapsed:.1f}", f"{grad_norm:.4f}"])
            csv_file.flush()
            t_log = now
            tokens_since_log = 0

        # --- Évaluation ---
        if step % cfg.eval_interval == 0 or step == cfg.max_steps:
            val_loss = estimate_loss(model, val_ds, cfg.batch_size,
                                     cfg.eval_batches, ctx)
            elapsed = elapsed_before + (time.perf_counter() - t_start)
            print(f"  >> eval step {step}: loss_val {val_loss:.4f} "
                  f"(train {loss_train_running:.4f})", flush=True)
            writer.writerow([step, tokens_seen, f"{loss_train_running:.6f}",
                             f"{val_loss:.6f}", f"{lr:.6e}", "", "",
                             f"{elapsed:.1f}", ""])
            csv_file.flush()

            if val_loss < best_val:
                best_val = val_loss
                save_checkpoint(best_path, raw_model, optimizer, step,
                                tokens_seen, cfg, best_val, elapsed)

            if cfg.overfit and loss_train_running < cfg.overfit_target:
                overfit_reached = True
                print(f"\n[surapprentissage] loss {loss_train_running:.4f} "
                      f"< {cfg.overfit_target} au step {step} — le test passe.")
                break

            # On remet à zéro les compteurs de débit après une évaluation.
            # Sinon l'intervalle de mesure suivant englobe le temps passé à
            # évaluer, et le débit rapporté s'effondre artificiellement — ce
            # qui donnait des lignes à 15 % de MFU alternant avec des lignes à
            # 55 %, sans que rien n'ait changé dans l'entraînement lui-même.
            t_log = time.perf_counter()
            tokens_since_log = 0

        # --- Instantané numéroté, pour la courbe de progression ---
        if (not cfg.overfit and cfg.snapshot_every
                and step % cfg.snapshot_every == 0):
            snap = os.path.join(cfg.out_dir,
                                f"{cfg.run_name}_step{step:06d}.pt")
            save_snapshot(snap, raw_model, step, tokens_seen, best_val)
            print(f"  [instantané] {snap} "
                  f"({os.path.getsize(snap)/1e6:.0f} Mo)", flush=True)

        # --- Checkpoint périodique, à l'horloge ---
        now = time.perf_counter()
        if not cfg.overfit and (now - t_last_ckpt) >= cfg.ckpt_minutes * 60:
            elapsed = elapsed_before + (now - t_start)
            save_checkpoint(ckpt_path, raw_model, optimizer, step, tokens_seen,
                            cfg, best_val, elapsed)
            t_last_ckpt = now
            print(f"  [checkpoint] step {step} -> {ckpt_path} "
                  f"({os.path.getsize(ckpt_path)/1e6:.0f} Mo)", flush=True)

    # --- Fin ---
    elapsed = elapsed_before + (time.perf_counter() - t_start)
    if not cfg.overfit:
        save_checkpoint(ckpt_path, raw_model, optimizer, step, tokens_seen,
                        cfg, best_val, elapsed)
    csv_file.close()

    summary = {
        "run_name": cfg.run_name,
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "mode": "surapprentissage" if cfg.overfit else "entrainement",
        "steps": step,
        "tokens_vus": tokens_seen,
        "loss_train_finale": loss_train_running,
        "best_val_loss": None if best_val == float("inf") else best_val,
        "duree_s": round(elapsed, 1),
        "parametres_non_embedding": n_params,
        "flops_par_token": flops_per_token,
        "peak_tflops_reference": MEASURED_PEAK_TFLOPS_BF16,
        "config": vars(cfg),
    }
    if cfg.overfit:
        summary["test_reussi"] = overfit_reached
        summary["cible_loss"] = cfg.overfit_target
        summary["parties"] = train_ds.n_games
        summary["tokens_du_jeu"] = train_ds.n_tokens

    suffix = "overfit" if cfg.overfit else "train"
    out = os.path.join(cfg.log_dir, f"{cfg.run_name}_{suffix}_summary.json")
    with open(out, "w") as f:
        json.dump(summary, f, indent=2, ensure_ascii=False)

    print(f"\n=== Terminé : {step} steps, {tokens_seen/1e6:.1f} M tokens, "
          f"{elapsed/60:.1f} min ===")
    print(f"résumé -> {out}")

    if cfg.overfit and not overfit_reached:
        raise SystemExit(
            f"\nECHEC : la loss n'est pas descendue sous {cfg.overfit_target} "
            f"en {cfg.overfit_steps} steps (elle vaut {loss_train_running:.4f}). "
            f"Il y a un bug — ne pas lancer le run complet.")


if __name__ == "__main__":
    main()
evaluate.py609 lines

Legality, human agreement, per-rule tests, complete games, Wilson intervals.

"""Phase 4 — Mesurer ce que le modèle a réellement appris.

Le chiffre central du projet est le **taux de coups légaux en génération
libre** : sur une position donnée, sans aucune contrainte imposée, quelle
proportion des coups que le modèle propose sont effectivement jouables ?

Ce chiffre mérite qu'on insiste sur ce qu'il mesure. Le modèle n'a jamais vu
d'échiquier, ne connaît pas les règles, et n'a aucun moyen de vérifier quoi que
ce soit. S'il propose un coup légal, c'est uniquement parce qu'il a inféré les
règles du jeu à partir de la seule régularité statistique de millions de
parties. Un taux élevé n'est donc pas une performance d'échecs : c'est la
preuve qu'un système entraîné à prédire des symboles a reconstruit une
mécanique sous-jacente qu'on ne lui a jamais décrite.

Toutes les mesures sont accompagnées d'un intervalle de confiance, parce qu'un
taux mesuré sur un échantillon fini est une estimation, pas une vérité.
"""

import argparse
import json
import math
import os
import random
import time
from collections import Counter

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import chess
import torch
import torch.nn.functional as F

from model import ChessGPT, ModelConfig


# ---------------------------------------------------------------------------
# Statistiques
# ---------------------------------------------------------------------------

def wilson_interval(successes: int, total: int, z: float = 1.96):
    """Intervalle de confiance de Wilson à 95 %.

    Pourquoi Wilson plutôt que l'intervalle « normal » enseigné partout
    (p ± 1.96 √(p(1-p)/n)) : ce dernier se comporte très mal quand le taux
    approche 0 ou 100 %. Il peut produire une borne supérieure au-dessus de
    100 %, ou un intervalle de largeur nulle si aucun échec n'est observé — ce
    qui affirmerait une certitude absolue à partir d'un échantillon fini.
    Comme on espère précisément mesurer des taux proches de 100 %, c'est le
    cas dégénéré qui nous concerne directement.
    """
    if total == 0:
        return (0.0, 0.0, 0.0)
    p = successes / total
    denom = 1 + z * z / total
    centre = (p + z * z / (2 * total)) / denom
    margin = (z * math.sqrt(p * (1 - p) / total + z * z / (4 * total * total))
              / denom)
    return p, max(0.0, centre - margin), min(1.0, centre + margin)


# ---------------------------------------------------------------------------
# Chargement
# ---------------------------------------------------------------------------

def load_model(ckpt_path: str, device: str):
    """Charge un checkpoint de reprise ou un instantané léger indifféremment.

    Les deux contiennent les poids, la configuration du modèle et les
    compteurs ; seul le checkpoint de reprise porte en plus l'état de
    l'optimiseur, dont l'évaluation n'a que faire.
    """
    ck = torch.load(ckpt_path, map_location=device, weights_only=False)
    cfg = ModelConfig(**ck["model_config"])
    model = ChessGPT(cfg).to(device)
    model.load_state_dict(ck["model"])
    model.eval()
    return model, ck


def load_vocab(path: str):
    with open(path) as f:
        v = json.load(f)
    return v["stoi"], v["itos"], v["bos_id"], v["eos_id"]


@torch.no_grad()
def last_logits_batched(model, prompts, device, batch_size, block_size):
    """Logits du dernier token, pour une liste de préfixes de longueurs variées.

    Subtilité qui coûte cher si on la rate. La façon naturelle de traiter des
    séquences de longueurs différentes en un seul batch est de les compléter
    par un token de remplissage. Mais notre modèle n'a pas de masque de
    remplissage : son attention est purement causale, donc chaque position
    regarde *tout* ce qui la précède, y compris le remplissage. Compléter à
    gauche avec des <bos> revient donc à faire croire au modèle que la partie a
    commencé plusieurs fois — ce qui pollue la prédiction.

    On regroupe donc les préfixes par longueur exacte : aucun remplissage,
    aucun token parasite, et on garde le bénéfice du traitement par lots.
    C'est d'autant plus important que le taux de coups légaux est le chiffre
    central du projet : une mesure biaisée par un artefact de remplissage
    n'aurait aucune valeur.
    """
    by_length = {}
    for i, p in enumerate(prompts):
        p = p[-block_size:]
        by_length.setdefault(len(p), []).append((i, p))

    out = [None] * len(prompts)
    for length, items in by_length.items():
        for start in range(0, len(items), batch_size):
            chunk = items[start:start + batch_size]
            batch = torch.tensor([p for _, p in chunk], dtype=torch.long,
                                 device=device)
            with torch.autocast("cuda", dtype=torch.bfloat16):
                logits, _ = model(batch)
            logits = logits[:, -1, :].float()
            for row, (idx, _) in enumerate(chunk):
                out[idx] = logits[row]
    return torch.stack(out)


# ---------------------------------------------------------------------------
# 1. Taux de coups légaux en génération libre
# ---------------------------------------------------------------------------

@torch.no_grad()
def legal_move_rate(model, stoi, itos, bos, games, device, n_positions,
                    temperature, batch_size, rng):
    """Sur des positions issues de vraies parties, un coup échantillonné
    librement est-il légal ?

    Protocole : on prend une partie du jeu de validation, on la rejoue jusqu'à
    un demi-coup tiré au hasard, on donne au modèle la séquence des coups
    joués, et on échantillonne **un** coup dans sa distribution de sortie sans
    aucun masquage. On vérifie ensuite la légalité avec python-chess.

    Aucune contrainte n'est appliquée : le modèle a le droit de proposer
    n'importe lequel des 1971 tokens du vocabulaire, y compris <pad>.
    """
    prompts, boards = [], []
    for _ in range(n_positions):
        game = rng.choice(games)
        moves = game.split()
        # On coupe entre le 1er et l'avant-dernier demi-coup
        cut = rng.randint(1, len(moves) - 1)
        board = chess.Board()
        for m in moves[:cut]:
            board.push(chess.Move.from_uci(m))
        prompts.append([bos] + [stoi[m] for m in moves[:cut]])
        boards.append(board)

    legal = 0
    illegal_reasons = Counter()

    all_logits = last_logits_batched(model, prompts, device, batch_size,
                                     model.cfg.block_size)

    if temperature == 0.0:
        ids = all_logits.argmax(dim=-1)
    else:
        probs = F.softmax(all_logits / temperature, dim=-1)
        ids = torch.multinomial(probs, num_samples=1).squeeze(-1)

    for i, tok_id in enumerate(ids.tolist()):
        token = itos[tok_id]
        if token in ("<pad>", "<bos>", "<eos>"):
            illegal_reasons["token_special"] += 1
            continue
        try:
            move = chess.Move.from_uci(token)
        except ValueError:
            illegal_reasons["uci_invalide"] += 1
            continue
        if move in boards[i].legal_moves:
            legal += 1
        else:
            # On distingue plusieurs échecs très différents : viser une case
            # vide, déplacer une pièce adverse, ou proposer un coup
            # géométriquement correct mais illégal dans cette position (roi
            # en échec, pièce clouée, case d'arrivée occupée par un allié).
            if boards[i].piece_at(move.from_square) is None:
                illegal_reasons["case_depart_vide"] += 1
            elif (boards[i].piece_at(move.from_square).color
                  != boards[i].turn):
                illegal_reasons["piece_adverse"] += 1
            else:
                illegal_reasons["deplacement_illegal"] += 1

    total = len(prompts)
    p, lo, hi = wilson_interval(legal, total)
    return {
        "positions_testees": total,
        "coups_legaux": legal,
        "taux": round(p, 5),
        "ic95_bas": round(lo, 5),
        "ic95_haut": round(hi, 5),
        "temperature": temperature,
        "motifs_echec": dict(illegal_reasons),
        "methode": ("un coup échantillonné sans masquage sur chaque position, "
                    "positions tirées à un demi-coup aléatoire de parties du "
                    "jeu de validation, légalité vérifiée par python-chess ; "
                    "intervalle de Wilson à 95 %"),
    }


# ---------------------------------------------------------------------------
# 2. Parties complètes sans aucun coup illégal
# ---------------------------------------------------------------------------

@torch.no_grad()
def full_game_rate(model, stoi, itos, bos, eos, device, n_games,
                   temperature, max_plies, rng, batch_size=64):
    """Le modèle joue des parties entières contre lui-même, depuis la position
    de départ, sans jamais être corrigé. Combien vont au bout sans faute ?

    C'est une mesure bien plus sévère que le taux par coup. Une partie de 80
    demi-coups n'est comptée comme réussie que si les 80 coups sont légaux.
    Même avec un taux par coup de 99 %, seules 45 % des parties passeraient
    (0.99^80). Ce chiffre teste donc autant la constance que la compétence.

    Les parties sont jouées **en parallèle**, par paquets. La version naïve —
    une partie après l'autre — enchaîne jusqu'à 250 passes avant successives
    avec une seule séquence à la fois, ce qui laisse le GPU à 20 % d'occupation :
    on paie le coût de lancement d'un noyau CUDA pour calculer un unique coup.
    En jouant 64 parties de front, la carte travaille sur 64 séquences à chaque
    passe et le temps d'évaluation s'effondre.

    Une subtilité rend ce parallélisme facile ici : toutes les parties du
    paquet démarrent au même moment et gagnent exactement un token par tour.
    Elles ont donc toujours la même longueur, et aucun remplissage n'est
    nécessaire. Les parties terminées avant les autres restent dans le tenseur
    — on ignore simplement ce qu'elles produisent.
    """
    successes = 0
    plies_before_error = []
    lengths = []
    terminations = Counter()
    block = model.cfg.block_size

    remaining = n_games
    while remaining > 0:
        b = min(batch_size, remaining)
        remaining -= b

        boards = [chess.Board() for _ in range(b)]
        seqs = torch.full((b, 1), bos, dtype=torch.long, device=device)
        active = [True] * b
        ok = [True] * b
        done_len = [0] * b

        for ply in range(max_plies):
            if not any(active):
                break
            ctx = seqs[:, -block:]
            with torch.autocast("cuda", dtype=torch.bfloat16):
                logits, _ = model(ctx)
            logits = logits[:, -1, :].float()

            if temperature == 0.0:
                ids = logits.argmax(dim=-1)
            else:
                probs = F.softmax(logits / temperature, dim=-1)
                ids = torch.multinomial(probs, num_samples=1).squeeze(-1)
            ids_list = ids.tolist()

            for i in range(b):
                if not active[i]:
                    continue
                token = itos[ids_list[i]]

                if token == "<eos>":
                    terminations["eos_predit"] += 1
                    active[i] = False
                    done_len[i] = ply
                    continue
                if token in ("<pad>", "<bos>"):
                    ok[i] = False
                    plies_before_error.append(ply)
                    terminations["token_special"] += 1
                    active[i] = False
                    done_len[i] = ply
                    continue
                try:
                    move = chess.Move.from_uci(token)
                except ValueError:
                    ok[i] = False
                    plies_before_error.append(ply)
                    terminations["uci_invalide"] += 1
                    active[i] = False
                    done_len[i] = ply
                    continue
                if move not in boards[i].legal_moves:
                    ok[i] = False
                    plies_before_error.append(ply)
                    terminations["coup_illegal"] += 1
                    active[i] = False
                    done_len[i] = ply
                    continue

                boards[i].push(move)
                if boards[i].is_game_over():
                    terminations["partie_terminee_regles"] += 1
                    active[i] = False
                    done_len[i] = ply + 1

            # Les parties terminées continuent de recevoir le token prédit :
            # leur contenu n'est plus lu, seule la forme du tenseur compte.
            seqs = torch.cat([seqs, ids.unsqueeze(1)], dim=1)

        for i in range(b):
            if active[i]:
                terminations["longueur_max_atteinte"] += 1
                done_len[i] = max_plies
            if ok[i]:
                successes += 1
            lengths.append(done_len[i])

    p, lo, hi = wilson_interval(successes, n_games)
    mean_err = (sum(plies_before_error) / len(plies_before_error)
                if plies_before_error else None)
    return {
        "parties_jouees": n_games,
        "parties_sans_coup_illegal": successes,
        "taux": round(p, 5),
        "ic95_bas": round(lo, 5),
        "ic95_haut": round(hi, 5),
        "temperature": temperature,
        "longueur_moyenne_demi_coups": round(sum(lengths) / len(lengths), 2),
        "demi_coups_moyens_avant_premiere_faute": (
            round(mean_err, 2) if mean_err is not None else "aucune faute"),
        "terminaisons": dict(terminations),
        "methode": ("parties générées depuis la position initiale, le modèle "
                    "joue les deux couleurs, aucun masquage, arrêt à la "
                    "première illégalité ou à <eos>"),
    }


# ---------------------------------------------------------------------------
# 3. Accord avec le coup humain
# ---------------------------------------------------------------------------

@torch.no_grad()
def human_agreement(model, stoi, itos, bos, games, device, n_positions,
                    batch_size, rng):
    """Le modèle propose-t-il le coup que l'humain a réellement joué ?

    Attention à l'interprétation : ce n'est pas une mesure de la qualité du
    jeu. Un moteur d'échecs parfait aurait un accord *faible* avec des joueurs
    de 2000 Elo, puisqu'il jouerait mieux qu'eux. Ce chiffre mesure la fidélité
    à la distribution d'entraînement, ce qui est exactement l'objectif d'un
    modèle de langage — mais ne dit rien de la force au jeu, que seule la
    calibration Elo de la phase 5 pourra établir.
    """
    prompts, targets = [], []
    for _ in range(n_positions):
        game = rng.choice(games)
        moves = game.split()
        cut = rng.randint(1, len(moves) - 1)
        prompts.append([bos] + [stoi[m] for m in moves[:cut]])
        targets.append(stoi[moves[cut]])

    top1 = top5 = 0
    all_logits = last_logits_batched(model, prompts, device, batch_size,
                                     model.cfg.block_size)
    top = all_logits.topk(5, dim=-1).indices.tolist()
    for i, t in enumerate(targets):
        if top[i][0] == t:
            top1 += 1
        if t in top[i]:
            top5 += 1

    p1, lo1, hi1 = wilson_interval(top1, n_positions)
    p5, lo5, hi5 = wilson_interval(top5, n_positions)
    return {
        "positions_testees": n_positions,
        "top1": {"taux": round(p1, 5), "ic95_bas": round(lo1, 5),
                 "ic95_haut": round(hi1, 5)},
        "top5": {"taux": round(p5, 5), "ic95_bas": round(lo5, 5),
                 "ic95_haut": round(hi5, 5)},
        "methode": ("comparaison du coup le plus probable (et des 5 plus "
                    "probables) au coup effectivement joué par l'humain, sur "
                    "des positions du jeu de validation ; intervalle de Wilson"),
    }


# ---------------------------------------------------------------------------
# 4. Tests par règle
# ---------------------------------------------------------------------------

def mine_rule_positions(games, stoi, bos, n_per_rule, rng, max_scan):
    """Extrait du corpus des positions illustrant chaque règle particulière.

    Plutôt que d'inventer des positions, on les prend dans de vraies parties :
    on cherche les moments où un humain a roqué, pris en passant, promu, ou
    joué alors qu'il était en échec. On dispose ainsi, pour chaque règle, de
    positions où le coup en question est non seulement légal mais pertinent.
    """
    buckets = {"roque": [], "en_passant": [], "promotion": [], "echec": []}
    scanned = 0

    for game in games:
        if scanned >= max_scan:
            break
        if all(len(v) >= n_per_rule for v in buckets.values()):
            break
        scanned += 1
        moves = game.split()
        board = chess.Board()
        prefix = [bos]
        for i, uci in enumerate(moves):
            move = chess.Move.from_uci(uci)
            entry = None
            if board.is_check() and len(buckets["echec"]) < n_per_rule:
                entry = ("echec", list(prefix), uci)
            elif (board.is_castling(move)
                  and len(buckets["roque"]) < n_per_rule):
                entry = ("roque", list(prefix), uci)
            elif (board.is_en_passant(move)
                  and len(buckets["en_passant"]) < n_per_rule):
                entry = ("en_passant", list(prefix), uci)
            elif move.promotion and len(buckets["promotion"]) < n_per_rule:
                entry = ("promotion", list(prefix), uci)

            if entry is not None and len(prefix) > 1:
                buckets[entry[0]].append(
                    {"prefix": entry[1], "human_move": entry[2],
                     "fen": board.fen()})

            prefix.append(stoi[uci])
            board.push(move)

    return buckets, scanned


@torch.no_grad()
def rule_tests(model, stoi, itos, bos, games, device, n_per_rule, rng,
               max_scan):
    """Pour chaque règle, mesure trois choses sur les positions extraites :

    - le coup le plus probable du modèle est-il légal ?
    - correspond-il au coup humain ?
    - la probabilité totale que le modèle place sur les coups légaux
      (une mesure plus fine que le simple top-1 : elle dit si le modèle a
      compris la contrainte ou s'il a seulement de la chance)
    """
    buckets, scanned = mine_rule_positions(games, stoi, bos, n_per_rule, rng,
                                           max_scan)
    results = {"_parties_parcourues": scanned}
    block = model.cfg.block_size

    for rule, entries in buckets.items():
        if not entries:
            results[rule] = {"positions": 0,
                             "note": "aucune position trouvée dans le corpus"}
            continue

        legal_top1 = 0
        match_human = 0
        legal_mass = []

        for e in entries:
            board = chess.Board(e["fen"])
            ids = e["prefix"][-block:]
            x = torch.tensor([ids], dtype=torch.long, device=device)
            with torch.autocast("cuda", dtype=torch.bfloat16):
                logits, _ = model(x)
            logits = logits[0, -1, :].float()
            probs = F.softmax(logits, dim=-1)

            legal_ids = [stoi[m.uci()] for m in board.legal_moves
                         if m.uci() in stoi]
            legal_mass.append(float(probs[legal_ids].sum()))

            best = itos[int(logits.argmax())]
            try:
                if chess.Move.from_uci(best) in board.legal_moves:
                    legal_top1 += 1
            except ValueError:
                pass
            if best == e["human_move"]:
                match_human += 1

        n = len(entries)
        pl, lol, hil = wilson_interval(legal_top1, n)
        pm, lom, him = wilson_interval(match_human, n)
        results[rule] = {
            "positions": n,
            "top1_legal": {"taux": round(pl, 5), "ic95_bas": round(lol, 5),
                           "ic95_haut": round(hil, 5)},
            "top1_egal_coup_humain": {"taux": round(pm, 5),
                                      "ic95_bas": round(lom, 5),
                                      "ic95_haut": round(him, 5)},
            "masse_de_probabilite_sur_coups_legaux": {
                "moyenne": round(sum(legal_mass) / n, 5),
                "min": round(min(legal_mass), 5),
                "max": round(max(legal_mass), 5),
            },
        }

    results["_methode"] = (
        "positions extraites de vraies parties du jeu de validation : pour "
        "'roque', 'en_passant' et 'promotion', les positions où l'humain a "
        "effectivement joué ce type de coup ; pour 'echec', les positions où "
        "le camp au trait est en échec et doit donc en sortir. Le modèle est "
        "interrogé sans masquage.")
    return results


# ---------------------------------------------------------------------------

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", required=True)
    p.add_argument("--data-dir", default="data")
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--val-games", default="data/val_games.txt",
                   help="parties de validation en clair, une par ligne")
    p.add_argument("--device", default="cuda:0",
                   help="cuda:0 = RTX 3060, pour évaluer pendant que le 3090 entraîne")
    p.add_argument("--out", default=None)
    p.add_argument("--seed", type=int, default=1234)

    p.add_argument("--n-legal", type=int, default=5000)
    p.add_argument("--n-full-games", type=int, default=200)
    p.add_argument("--n-agreement", type=int, default=5000)
    p.add_argument("--n-per-rule", type=int, default=300)
    p.add_argument("--max-scan", type=int, default=200000)
    p.add_argument("--temperature", type=float, default=1.0)
    p.add_argument("--max-plies", type=int, default=250)
    p.add_argument("--batch-size", type=int, default=128)
    p.add_argument("--skip-full-games", action="store_true")

    args = p.parse_args()
    rng = random.Random(args.seed)
    torch.manual_seed(args.seed)

    stoi, itos, bos, eos = load_vocab(args.vocab)
    model, ck = load_model(args.ckpt, args.device)
    print(f"[modèle] {args.ckpt} | step {ck['step']:,} | "
          f"{ck['tokens_seen']:,} tokens vus")

    with open(args.val_games) as f:
        games = [l.strip() for l in f if l.strip().count(" ") >= 4]
    print(f"[données] {len(games):,} parties de validation")

    report = {
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "checkpoint": os.path.abspath(args.ckpt),
        "step": ck["step"],
        "tokens_vus": ck["tokens_seen"],
        "seed": args.seed,
        "parties_validation_disponibles": len(games),
    }

    t0 = time.perf_counter()
    print("\n[1/4] Taux de coups légaux en génération libre...", flush=True)
    report["legalite_generation_libre"] = legal_move_rate(
        model, stoi, itos, bos, games, args.device, args.n_legal,
        args.temperature, args.batch_size, rng)
    r = report["legalite_generation_libre"]
    print(f"      {r['taux']:.2%} "
          f"[IC95 {r['ic95_bas']:.2%} - {r['ic95_haut']:.2%}] "
          f"sur {r['positions_testees']:,} positions")

    print("\n[2/4] Accord avec le coup humain...", flush=True)
    report["accord_humain"] = human_agreement(
        model, stoi, itos, bos, games, args.device, args.n_agreement,
        args.batch_size, rng)
    r = report["accord_humain"]
    print(f"      top-1 {r['top1']['taux']:.2%} | top-5 {r['top5']['taux']:.2%}")

    print("\n[3/4] Tests par règle...", flush=True)
    report["tests_par_regle"] = rule_tests(
        model, stoi, itos, bos, games, args.device, args.n_per_rule, rng,
        args.max_scan)
    for rule in ("roque", "en_passant", "promotion", "echec"):
        d = report["tests_par_regle"].get(rule, {})
        if d.get("positions"):
            print(f"      {rule:<12} n={d['positions']:>4} | "
                  f"top1 légal {d['top1_legal']['taux']:.2%} | "
                  f"= humain {d['top1_egal_coup_humain']['taux']:.2%} | "
                  f"masse légale {d['masse_de_probabilite_sur_coups_legaux']['moyenne']:.2%}")

    if args.skip_full_games:
        report["parties_completes"] = "non mesuré (--skip-full-games)"
        print("\n[4/4] Parties complètes : ignoré")
    else:
        print("\n[4/4] Parties complètes sans coup illégal...", flush=True)
        report["parties_completes"] = full_game_rate(
            model, stoi, itos, bos, eos, args.device, args.n_full_games,
            args.temperature, args.max_plies, rng, args.batch_size)
        r = report["parties_completes"]
        print(f"      {r['taux']:.2%} "
              f"[IC95 {r['ic95_bas']:.2%} - {r['ic95_haut']:.2%}] "
              f"sur {r['parties_jouees']} parties")

    report["duree_evaluation_s"] = round(time.perf_counter() - t0, 1)

    out = args.out or os.path.join(
        "logs", f"eval_step{ck['step']}.json")
    os.makedirs(os.path.dirname(out), exist_ok=True)
    with open(out, "w") as f:
        json.dump(report, f, indent=2, ensure_ascii=False)
    print(f"\nRapport -> {out} ({report['duree_evaluation_s']:.0f} s)")


if __name__ == "__main__":
    main()
eval_watcher.py147 lines

Evaluates the snapshots as training goes along, on the second card, in subprocesses.

"""Phase 3 — Évaluer les instantanés au fil de l'entraînement, sur l'autre carte.

Pendant que le 3090 entraîne, ce script surveille l'apparition d'instantanés
numérotés et les évalue sur le 3060. On obtient ainsi, sans ralentir
l'entraînement d'une seconde, la courbe qui montre le modèle découvrir les
règles du jeu : le taux de coups légaux en fonction du nombre de tokens vus.

C'est le graphique le plus parlant du projet, et il a une particularité — on ne
peut pas le reconstituer après coup. Si on n'évalue pas les états
intermédiaires pendant qu'ils existent, ils sont perdus. D'où ce script, lancé
en parallèle du run.

Usage :
    nohup python eval_watcher.py --run-name run1 > logs/eval_watcher.log 2>&1 &
"""

import argparse
import glob
import json
import os
import re
import subprocess
import sys
import time

STEP_RE = re.compile(r"_step(\d+)\.pt$")


def find_snapshots(ckpt_dir, run_name):
    pattern = os.path.join(ckpt_dir, f"{run_name}_step*.pt")
    out = []
    for path in glob.glob(pattern):
        m = STEP_RE.search(path)
        if m:
            out.append((int(m.group(1)), path))
    return sorted(out)


def already_done(log_dir, step, prefix="eval"):
    return os.path.exists(os.path.join(log_dir, f"{prefix}_step{step}.json"))


def evaluate_one(python, step, path, args):
    """Lance evaluate.py en sous-processus.

    Pourquoi un sous-processus plutôt qu'un import : chaque évaluation charge
    un modèle sur le GPU, et PyTorch ne rend jamais complètement la mémoire
    d'un modèle libéré au sein d'un même processus. Sur vingt évaluations
    successives, la fragmentation finirait par provoquer un dépassement
    mémoire. Un processus qui meurt rend tout, sans exception.
    """
    cmd = [
        python, "evaluate.py",
        "--ckpt", path,
        "--device", args.device,
        "--vocab", args.vocab,
        "--val-games", args.val_games,
        "--n-legal", str(args.n_legal),
        "--n-agreement", str(args.n_agreement),
        "--n-per-rule", str(args.n_per_rule),
        "--n-full-games", str(args.n_full_games),
        "--out", os.path.join(args.log_dir, f"{args.prefix}_step{step}.json"),
        "--temperature", str(args.temperature),
    ]
    print(f"[{time.strftime('%H:%M:%S')}] évaluation du step {step}...",
          flush=True)
    t0 = time.perf_counter()
    r = subprocess.run(cmd, capture_output=True, text=True)
    if r.returncode != 0:
        print(f"  ECHEC (code {r.returncode})", flush=True)
        print(r.stdout[-2000:], flush=True)
        print(r.stderr[-2000:], flush=True)
        return False
    # On ne réaffiche que la ligne du chiffre central.
    for line in r.stdout.splitlines():
        if "%" in line and "[IC95" in line:
            print(f"  {line.strip()}", flush=True)
            break
    print(f"  terminé en {time.perf_counter()-t0:.0f} s", flush=True)
    return True


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--run-name", default="run1")
    p.add_argument("--ckpt-dir", default="checkpoints")
    p.add_argument("--log-dir", default="logs")
    p.add_argument("--device", default="cuda:0", help="cuda:0 = RTX 3060")
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--val-games", default="data/val_games.txt")
    p.add_argument("--python", default=sys.executable)
    p.add_argument("--poll-seconds", type=int, default=120)
    # Sans préfixe distinct, les rapports de deux runs différents s'écrasent
    # mutuellement dans logs/. C'est arrivé de justesse entre run1 et run2.
    p.add_argument("--prefix", default="eval",
                   help="préfixe des rapports, ex. eval_run2")
    p.add_argument("--temperature", type=float, default=1.0,
                   help="température de génération pour les parties complètes")
    # Échantillons volontairement réduits par rapport à l'évaluation finale :
    # ici on veut une courbe, pas un chiffre publiable. L'évaluation finale
    # du meilleur checkpoint utilisera les tailles par défaut, bien plus
    # grandes, pour resserrer les intervalles de confiance.
    p.add_argument("--n-legal", type=int, default=2000)
    p.add_argument("--n-agreement", type=int, default=2000)
    p.add_argument("--n-per-rule", type=int, default=150)
    p.add_argument("--n-full-games", type=int, default=60)
    # 150 min et non 45 : ce délai doit dépasser confortablement l'intervalle
    # entre deux instantanés, sinon le watcher expire entre deux et s'arrête
    # en plein run. Sur run2, un instantané tous les 5000 steps à 50 000
    # tokens/s fait 68 minutes — le réglage initial de 45 min garantissait
    # l'arrêt prématuré. Vérifier ce rapport à chaque changement de cadence
    # d'instantanés ou de débit.
    p.add_argument("--stop-when-idle-minutes", type=float, default=150.0,
                   help="s'arrête si aucun nouvel instantané pendant ce délai ; "
                        "doit dépasser l'intervalle entre deux instantanés")
    args = p.parse_args()

    print(f"Surveillance de {args.ckpt_dir}/{args.run_name}_step*.pt "
          f"sur {args.device}", flush=True)
    last_activity = time.perf_counter()

    while True:
        snapshots = find_snapshots(args.ckpt_dir, args.run_name)
        pending = [(s, p_) for s, p_ in snapshots
                   if not already_done(args.log_dir, s, args.prefix)]

        if pending:
            for step, path in pending:
                # On attend que le fichier soit stable : l'écriture est
                # atomique (rename), mais mieux vaut ne pas courir après.
                size = os.path.getsize(path)
                time.sleep(2)
                if os.path.getsize(path) != size:
                    continue
                evaluate_one(args.python, step, path, args)
                last_activity = time.perf_counter()
        else:
            idle_min = (time.perf_counter() - last_activity) / 60
            if idle_min >= args.stop_when_idle_minutes:
                print(f"Aucun nouvel instantané depuis {idle_min:.0f} min — "
                      f"arrêt de la surveillance.", flush=True)
                break
            time.sleep(args.poll_seconds)


if __name__ == "__main__":
    main()
engine.py246 lines

UCI engine with a legality mask, usable in any chess interface.

"""Un moteur d'échecs UCI standard, autour du modèle.

Deux usages :

    # comme moteur en ligne de commande (protocole UCI, pour cutechess/lichess-bot)
    python engine.py --ckpt checkpoints/run1_best.pt

    # comme bibliothèque
    from engine import ChessEngine
    eng = ChessEngine("checkpoints/run1_best.pt", device="cuda:0")
    move = eng.best_move(board)

Le point important est le **masque de légalité**. Le modèle, laissé libre,
propose parfois un coup impossible — c'est précisément ce que mesure la phase
4. Mais pour *jouer*, un tel coup est disqualifiant : un moteur UCI qui répond
un coup illégal perd la partie sur-le-champ.

La solution retenue est la plus simple possible : un seul passage avant, puis
on met à moins l'infini les logits de tous les tokens qui ne correspondent pas
à un coup légal dans la position courante, avant de prendre le maximum. Le
modèle choisit donc toujours parmi les coups légaux, en conservant l'ordre de
préférence qu'il avait appris.

On aurait pu faire autrement — rééchantillonner jusqu'à tomber sur un coup
légal, ou faire une recherche arborescente. Le masque a été préféré parce
qu'il est déterministe, qu'il coûte exactement un passage avant, et qu'il
n'altère pas les préférences relatives du modèle entre coups légaux. C'était
aussi le critère annoncé : à qualité comparable, on choisit ce qui s'explique
le plus simplement.
"""

import argparse
import json
import os
import sys

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import chess
import torch
import torch.nn.functional as F

from model import ChessGPT, ModelConfig


class ChessEngine:

    def __init__(self, ckpt_path: str, vocab_path: str = "data/vocab.json",
                 device: str = "cuda:0", temperature: float = 0.0,
                 top_k: int | None = None):
        self.device = device
        self.temperature = temperature
        self.top_k = top_k

        with open(vocab_path) as f:
            v = json.load(f)
        self.stoi = v["stoi"]
        self.itos = v["itos"]
        self.bos = v["bos_id"]
        self.eos = v["eos_id"]
        self.vocab_size = v["vocab_size"]

        ck = torch.load(ckpt_path, map_location=device, weights_only=False)
        cfg = ModelConfig(**ck["model_config"])
        self.model = ChessGPT(cfg).to(device)
        self.model.load_state_dict(ck["model"])
        self.model.eval()
        self.block_size = cfg.block_size
        self.ckpt_step = ck.get("step", 0)
        self.tokens_seen = ck.get("tokens_seen", 0)

        # Le modèle ne voit que la liste des coups depuis le début de la
        # partie. On la maintient ici en parallèle de l'échiquier.
        self.history: list[int] = []

    # -- gestion de la position ---------------------------------------------

    def set_position(self, board: chess.Board, move_history: list[str]):
        """Renseigne la suite de coups joués depuis la position initiale.

        Le modèle n'a aucune représentation de l'échiquier : il ne comprend
        qu'une séquence de coups. Une position transmise sous forme de FEN sans
        historique lui est donc inutilisable — d'où la limitation documentée
        plus bas dans `uci_loop`.
        """
        self.history = [self.stoi[m] for m in move_history if m in self.stoi]

    # -- le cœur -------------------------------------------------------------

    @torch.no_grad()
    def move_scores(self, board: chess.Board):
        """Renvoie {coup UCI: probabilité} restreint aux coups légaux."""
        ids = [self.bos] + self.history
        ids = ids[-self.block_size:]
        x = torch.tensor([ids], dtype=torch.long, device=self.device)

        with torch.autocast("cuda", dtype=torch.bfloat16) \
                if self.device.startswith("cuda") else _null():
            logits, _ = self.model(x)
        logits = logits[0, -1, :].float()

        # Le masque : un booléen par entrée du vocabulaire.
        mask = torch.zeros(self.vocab_size, dtype=torch.bool,
                           device=self.device)
        legal_moves = list(board.legal_moves)
        for move in legal_moves:
            uci = move.uci()
            idx = self.stoi.get(uci)
            if idx is not None:
                mask[idx] = True

        if not mask.any():
            # Ne devrait pas arriver : le vocabulaire est un sur-ensemble des
            # coups légaux. Si ça arrive, mieux vaut un coup légal au hasard
            # qu'un plantage en pleine partie.
            return {m.uci(): 1.0 / len(legal_moves) for m in legal_moves}

        logits = logits.masked_fill(~mask, float("-inf"))
        probs = F.softmax(logits, dim=-1)
        return {self.itos[i]: float(probs[i])
                for i in torch.nonzero(mask).flatten().tolist()}

    @torch.no_grad()
    def best_move(self, board: chess.Board) -> chess.Move:
        scores = self.move_scores(board)
        if self.temperature == 0.0:
            uci = max(scores, key=scores.get)
        else:
            items = list(scores.items())
            logits = torch.tensor([s for _, s in items]).log() / self.temperature
            if self.top_k:
                k = min(self.top_k, len(items))
                keep = logits.topk(k).indices
                filt = torch.full_like(logits, float("-inf"))
                filt[keep] = logits[keep]
                logits = filt
            i = int(torch.multinomial(F.softmax(logits, dim=-1), 1))
            uci = items[i][0]
        return chess.Move.from_uci(uci)

    # -- protocole UCI -------------------------------------------------------

    def uci_loop(self):
        """Boucle du protocole UCI, lue sur stdin, réponses sur stdout.

        Limitation assumée et importante : la commande `position fen ...` n'est
        pas réellement supportée, parce que le modèle a besoin de l'historique
        des coups et non d'une position. On accepte `position startpos moves
        ...`, qui est ce qu'envoient cutechess-cli et lichess-bot en jeu
        normal. Sur un `position fen`, le moteur joue quand même un coup légal
        — il se rabat sur le coup le plus probable *sans historique*, ce qui
        revient à jouer comme en début de partie. Le comportement est dégradé
        mais jamais illégal.
        """
        board = chess.Board()
        history: list[str] = []

        for line in sys.stdin:
            line = line.strip()
            if not line:
                continue
            parts = line.split()
            cmd = parts[0]

            if cmd == "uci":
                print(f"id name ChessGPT-{self.ckpt_step}")
                print("id author billy + claude")
                print("uciok", flush=True)

            elif cmd == "isready":
                print("readyok", flush=True)

            elif cmd == "ucinewgame":
                board = chess.Board()
                history = []
                self.history = []

            elif cmd == "position":
                if "startpos" in parts:
                    board = chess.Board()
                    history = []
                    if "moves" in parts:
                        for uci in parts[parts.index("moves") + 1:]:
                            board.push(chess.Move.from_uci(uci))
                            history.append(uci)
                elif "fen" in parts:
                    i = parts.index("fen")
                    end = parts.index("moves") if "moves" in parts else len(parts)
                    board = chess.Board(" ".join(parts[i + 1:end]))
                    history = []
                    if "moves" in parts:
                        for uci in parts[parts.index("moves") + 1:]:
                            board.push(chess.Move.from_uci(uci))
                            history.append(uci)
                self.set_position(board, history)

            elif cmd == "go":
                move = self.best_move(board)
                print(f"bestmove {move.uci()}", flush=True)

            elif cmd in ("quit", "stop"):
                if cmd == "quit":
                    break

            elif cmd == "setoption":
                pass   # aucune option exposée pour l'instant


class _null:
    def __enter__(self): return None
    def __exit__(self, *a): return False


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", required=True)
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--device", default="cuda:0")
    p.add_argument("--temperature", type=float, default=0.0)
    p.add_argument("--top-k", type=int, default=None)
    p.add_argument("--selftest", action="store_true",
                   help="joue une partie contre soi-même et vérifie la légalité")
    args = p.parse_args()

    eng = ChessEngine(args.ckpt, args.vocab, args.device,
                      args.temperature, args.top_k)

    if args.selftest:
        board = chess.Board()
        history = []
        while not board.is_game_over() and len(history) < 300:
            eng.set_position(board, history)
            move = eng.best_move(board)
            assert move in board.legal_moves, f"coup illégal produit : {move}"
            board.push(move)
            history.append(move.uci())
        print(f"partie complète : {len(history)} demi-coups, "
              f"résultat {board.result()}, aucun coup illégal")
        print(" ".join(history[:40]) + (" ..." if len(history) > 40 else ""))
        return

    eng.uci_loop()


if __name__ == "__main__":
    main()
engine_client.py89 lines

Lightweight engine client, which talks to the shared inference server.

"""Moteur UCI ultra-léger : NE charge pas le modèle, relaie vers infer_server.

lichess-bot lance un moteur par partie. Pour un modèle sans état, charger une
copie du modèle par partie est un gaspillage : ce client ne fait que parler le
protocole UCI et transmettre l'historique de coups au serveur d'inférence
(infer_server.py), qui détient l'unique copie du modèle.

Conséquences :
  - démarrage **instantané** (aucun import de torch, aucun chargement de poids) ;
  - empreinte minime (~quelques dizaines de Mo) ;
  - la concurrence n'est plus limitée par la mémoire mais par lichess-bot et
    Lichess. 50 clients légers + 1 modèle, au lieu de 50 modèles.

Le serveur applique le masque de légalité, donc ce client ne renvoie jamais un
coup illégal. Socket configurable via la variable d'environnement INFER_SOCKET.

Limitation identique à engine.py : seul `position startpos moves ...` est
géré (le modèle a besoin de l'historique, pas d'une FEN).
"""

import json
import os
import socket
import sys
import time

SOCK = os.environ.get("INFER_SOCKET",
                      "/mnt/ia-data/fine-tuning/chess/logs/infer.sock")


def demander(moves, tentatives=40):
    """Envoie l'historique au serveur, renvoie le coup. Réessaie si le serveur
    redémarre (le watchdog peut le relancer sous nous)."""
    dernier = None
    for _ in range(tentatives):
        try:
            s = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
            s.settimeout(30)
            s.connect(SOCK)
            s.sendall((json.dumps({"moves": moves}) + "\n").encode())
            data = b""
            while not data.endswith(b"\n"):
                chunk = s.recv(4096)
                if not chunk:
                    break
                data += chunk
            s.close()
            resp = json.loads(data)
            if resp.get("move"):
                return resp["move"]
            dernier = resp.get("error", "réponse vide")
        except (FileNotFoundError, ConnectionRefusedError, OSError) as e:
            dernier = str(e)
            time.sleep(0.5)   # serveur peut-être en train de (re)démarrer
    raise RuntimeError(f"serveur d'inférence injoignable : {dernier}")


def main():
    moves = []
    for line in sys.stdin:
        line = line.strip()
        if not line:
            continue
        parts = line.split()
        cmd = parts[0]

        if cmd == "uci":
            print("id name philidor-142M")
            print("id author billy + claude")
            print("uciok", flush=True)
        elif cmd == "isready":
            print("readyok", flush=True)
        elif cmd == "ucinewgame":
            moves = []
        elif cmd == "position":
            if "startpos" in parts and "moves" in parts:
                moves = parts[parts.index("moves") + 1:]
            elif "startpos" in parts:
                moves = []
            # `position fen ...` non supporté : le modèle a besoin de l'historique
        elif cmd == "go":
            print(f"bestmove {demander(moves)}", flush=True)
        elif cmd in ("quit", "stop"):
            if cmd == "quit":
                break


if __name__ == "__main__":
    main()
infer_server.py97 lines

Inference server: a single model in memory, several clients.

"""Serveur d'inférence : LE modèle, chargé une seule fois, pour N parties.

Un modèle sans état n'a aucune raison d'être dupliqué par partie. Ce serveur
charge le modèle une fois et répond aux clients UCI légers (engine_client.py)
via une socket Unix locale. Chaque requête est un historique de coups depuis la
position initiale ; la réponse est le meilleur coup **légal** (le masque de
légalité est appliqué ici, comme dans engine.py).

    python infer_server.py --ckpt checkpoints/run2_best.pt --device cuda:0 \
        --socket /mnt/ia-data/fine-tuning/chess/logs/infer.sock

Concurrence : chaque connexion est servie dans un thread, mais un verrou
sérialise les passes avant (une à la fois). Ce n'est pas un goulot : une passe
sur ce modèle coûte quelques millisecondes, largement de quoi alimenter des
dizaines de parties en rapid. Le coût mémoire, lui, reste celui d'UN modèle.

Protocole, une ligne JSON par sens :
    client -> serveur : {"moves": ["e2e4", "e7e5", ...]}
    serveur -> client : {"move": "g1f3"}   ou   {"error": "..."}
"""

import argparse
import json
import os
import socket
import socketserver
import threading

import chess

from engine import ChessEngine


class Handler(socketserver.StreamRequestHandler):
    def handle(self):
        for raw in self.rfile:
            try:
                req = json.loads(raw)
                moves = req.get("moves", [])
                board = chess.Board()
                for m in moves:
                    board.push(chess.Move.from_uci(m))
                with self.server.lock:
                    self.server.engine.set_position(board, moves)
                    mv = self.server.engine.best_move(board)
                resp = {"move": mv.uci()}
            except Exception as e:  # jamais planter le serveur pour une requête
                resp = {"error": str(e)}
            try:
                self.wfile.write((json.dumps(resp) + "\n").encode())
                self.wfile.flush()
            except (BrokenPipeError, ConnectionResetError):
                break


class Server(socketserver.ThreadingUnixStreamServer):
    daemon_threads = True
    allow_reuse_address = True


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", required=True)
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--device", default="cuda:0")
    p.add_argument("--temperature", type=float, default=0.0)
    p.add_argument("--socket", required=True)
    args = p.parse_args()

    ready = args.socket + ".ready"
    for path in (args.socket, ready):
        if os.path.exists(path):
            os.unlink(path)

    eng = ChessEngine(args.ckpt, args.vocab, args.device, args.temperature)

    srv = Server(args.socket, Handler)
    srv.engine = eng
    srv.lock = threading.Lock()   # une passe avant à la fois
    os.chmod(args.socket, 0o600)

    # Signal de disponibilité : le watchdog attend ce fichier avant de lancer
    # lichess-bot, sinon les premiers clients trouvent une socket morte.
    with open(ready, "w") as f:
        f.write("ok\n")
    print(f"infer_server prêt sur {args.socket} (device={args.device})", flush=True)

    try:
        srv.serve_forever()
    finally:
        for path in (args.socket, ready):
            if os.path.exists(path):
                os.unlink(path)


if __name__ == "__main__":
    main()
elo_match.py489 lines

Matches against Stockfish, ladder between snapshots, duels between two models.

"""Phase 5 — Calibration de la force de jeu.

On fait jouer le modèle contre Stockfish bridé à différents niveaux, puis on
convertit le score obtenu en différence d'Elo.

### Pourquoi pas cutechess-cli

Le plan initial prévoyait `cutechess-cli`, l'outil standard pour organiser des
tournois entre moteurs. Il n'est empaqueté ni dans Homebrew ni dans les dépôts
apt de cette machine, et le compiler demande la chaîne Qt6 complète, donc les
droits administrateur. Le juge de match tient en deux cents lignes de
python-chess, qui est déjà une dépendance du projet : on le fait à la main.
C'était aussi le critère annoncé au départ — à qualité comparable, choisir ce
qui s'explique le plus simplement.

Le seul renoncement est l'ouverture imposée : cutechess sait distribuer un
livre d'ouvertures pour diversifier les parties. On compense en jouant les
premiers coups au hasard parmi les coups légaux (voir `--random-plies`), ce qui
évite que les deux moteurs rejouent deux cents fois la même partie.

### Comment on convertit un score en Elo

La formule d'Elo dit qu'un joueur ayant D points d'écart avec son adversaire
obtient en moyenne un score de 1 / (1 + 10^(-D/400)). On l'inverse : à partir
du score mesuré, on remonte à l'écart.

    D = -400 x log10(1/score - 1)

Deux précautions. D'abord, un score de 0 ou de 100 % donne un écart infini :
on ne peut alors que donner une borne. Ensuite, le score est mesuré sur un
échantillon fini, donc il a une incertitude — qu'on propage jusqu'à l'Elo au
lieu d'annoncer un chiffre net qui n'existe pas.
"""

import argparse
import json
import math
import os
import random
import time
from collections import Counter

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import chess
import chess.engine

from engine import ChessEngine


# ---------------------------------------------------------------------------
# Conversion score -> Elo
# ---------------------------------------------------------------------------

def score_to_elo(score: float):
    """Écart d'Elo correspondant à un score entre 0 et 1."""
    if score <= 0.0:
        return float("-inf")
    if score >= 1.0:
        return float("inf")
    return -400.0 * math.log10(1.0 / score - 1.0)


def elo_with_error(wins: int, draws: int, losses: int, z: float = 1.96):
    """Écart d'Elo et intervalle de confiance à 95 %.

    L'incertitude vient de l'échantillonnage : sur 200 parties, un score de
    60 % ne signifie pas que la vraie valeur est 60.0 %. On estime l'écart-type
    du score à partir de la distribution observée des résultats (victoire = 1,
    nulle = 0.5, défaite = 0), puis on convertit les deux bornes du score en
    bornes d'Elo.

    Les nulles comptent double dans la précision : elles réduisent la variance
    du score, donc resserrent l'intervalle. C'est pourquoi on ne peut pas
    utiliser une simple loi binomiale ici — il faut la variance réelle des
    trois issues.
    """
    n = wins + draws + losses
    if n == 0:
        return None

    score = (wins + 0.5 * draws) / n
    # Variance des résultats individuels autour du score moyen
    variance = (wins * (1.0 - score) ** 2
                + draws * (0.5 - score) ** 2
                + losses * (0.0 - score) ** 2) / n
    stderr = math.sqrt(variance / n) if n > 0 else 0.0

    lo_score = max(0.0, score - z * stderr)
    hi_score = min(1.0, score + z * stderr)

    elo = score_to_elo(score)
    elo_lo = score_to_elo(lo_score)
    elo_hi = score_to_elo(hi_score)

    def fmt(v):
        return None if math.isinf(v) else round(v, 1)

    return {
        "parties": n,
        "victoires": wins,
        "nulles": draws,
        "defaites": losses,
        "score": round(score, 4),
        "erreur_type_score": round(stderr, 5),
        "elo_diff": fmt(elo),
        "elo_diff_ic95_bas": fmt(elo_lo),
        "elo_diff_ic95_haut": fmt(elo_hi),
        "marge_erreur": (None if (math.isinf(elo) or math.isinf(elo_lo)
                                  or math.isinf(elo_hi))
                         else round((elo_hi - elo_lo) / 2, 1)),
    }


# ---------------------------------------------------------------------------
# Une partie
# ---------------------------------------------------------------------------

def play_game(model_engine, sf, model_is_white: bool, sf_limit,
              random_plies: int, rng, max_plies: int = 400):
    """Joue une partie. Renvoie (résultat du point de vue du modèle, raison).

    Résultat : 1.0 victoire, 0.5 nulle, 0.0 défaite.
    """
    board = chess.Board()
    history: list[str] = []

    # Quelques coups au hasard pour diversifier les parties, joués par les
    # deux camps. Sans ça, deux moteurs déterministes rejouent exactement la
    # même partie à chaque fois et l'échantillon n'a aucune valeur.
    for _ in range(random_plies):
        if board.is_game_over():
            break
        move = rng.choice(list(board.legal_moves))
        board.push(move)
        history.append(move.uci())

    while not board.is_game_over(claim_draw=True) and len(history) < max_plies:
        model_turn = (board.turn == chess.WHITE) == model_is_white
        if model_turn:
            model_engine.set_position(board, history)
            move = model_engine.best_move(board)
            if move not in board.legal_moves:
                # Ne devrait jamais arriver : le masque de légalité l'interdit.
                return (0.0 if model_is_white else 0.0), "coup_illegal_du_modele"
        else:
            move = sf.play(board, sf_limit).move
            if move is None:
                break
        board.push(move)
        history.append(move.uci())

    if len(history) >= max_plies:
        return 0.5, "limite_de_coups"

    outcome = board.outcome(claim_draw=True)
    if outcome is None:
        return 0.5, "indetermine"
    if outcome.winner is None:
        return 0.5, outcome.termination.name.lower()
    model_won = (outcome.winner == chess.WHITE) == model_is_white
    return (1.0 if model_won else 0.0), outcome.termination.name.lower()


def play_game_model_vs_model(eng_a, eng_b, a_is_white, random_plies, rng,
                             max_plies=400):
    """Deux checkpoints du même modèle s'affrontent.

    Renvoie le résultat du point de vue de `eng_a`. Sert à mesurer la
    progression : chaque instantané joue contre le modèle final, et l'écart de
    score se convertit en écart d'Elo. On obtient ainsi une courbe de force en
    fonction des tokens vus, sans avoir besoin d'un adversaire externe — donc
    sans dépendre du niveau exact auquel Stockfish est bridé.
    """
    board = chess.Board()
    history: list[str] = []

    for _ in range(random_plies):
        if board.is_game_over():
            break
        move = rng.choice(list(board.legal_moves))
        board.push(move)
        history.append(move.uci())

    while not board.is_game_over(claim_draw=True) and len(history) < max_plies:
        a_turn = (board.turn == chess.WHITE) == a_is_white
        eng = eng_a if a_turn else eng_b
        eng.set_position(board, history)
        move = eng.best_move(board)
        board.push(move)
        history.append(move.uci())

    if len(history) >= max_plies:
        return 0.5, "limite_de_coups"
    outcome = board.outcome(claim_draw=True)
    if outcome is None:
        return 0.5, "indetermine"
    if outcome.winner is None:
        return 0.5, outcome.termination.name.lower()
    a_won = (outcome.winner == chess.WHITE) == a_is_white
    return (1.0 if a_won else 0.0), outcome.termination.name.lower()


def run_head_to_head(ckpt_a, ckpt_b, vocab, device, n_games, random_plies,
                     seed, temperature, label_a="A", label_b="B"):
    """Duel direct entre deux checkpoints, avec assez de parties pour trancher.

    Les deux modèles peuvent avoir des architectures différentes : chaque
    checkpoint porte sa propre configuration, et `ChessEngine` la lit. On peut
    donc faire s'affronter un modèle de 51 M et un de 142 M sans rien changer.

    Note sur le déterminisme : à température 0, les deux moteurs jouent toujours
    le même coup dans une position donnée. Une partie est donc entièrement
    déterminée par ses coups d'ouverture aléatoires. C'est voulu — c'est
    l'équivalent du livre d'ouvertures qu'utilisent les tournois de moteurs, et
    ça garantit que la différence mesurée vient du jeu et non du hasard
    d'échantillonnage.
    """
    eng_a = ChessEngine(ckpt_a, vocab, device, temperature)
    eng_b = ChessEngine(ckpt_b, vocab, device, temperature)
    rng = random.Random(seed)
    wins = draws = losses = 0
    reasons = Counter()
    t0 = time.perf_counter()

    for i in range(n_games):
        r, reason = play_game_model_vs_model(eng_a, eng_b, i % 2 == 0,
                                             random_plies, rng)
        reasons[reason] += 1
        if r == 1.0:
            wins += 1
        elif r == 0.5:
            draws += 1
        else:
            losses += 1
        if (i + 1) % 25 == 0:
            sc = (wins + 0.5 * draws) / (i + 1)
            print(f"    {i+1}/{n_games} | +{wins} ={draws} -{losses} | "
                  f"score {sc:.1%} | {time.perf_counter()-t0:.0f} s", flush=True)

    stats = elo_with_error(wins, draws, losses)
    stats["modele_a"] = {"label": label_a, "checkpoint": os.path.abspath(ckpt_a),
                         "step": eng_a.ckpt_step, "tokens_vus": eng_a.tokens_seen}
    stats["modele_b"] = {"label": label_b, "checkpoint": os.path.abspath(ckpt_b),
                         "step": eng_b.ckpt_step, "tokens_vus": eng_b.tokens_seen}
    stats["duree_s"] = round(time.perf_counter() - t0, 1)
    stats["terminaisons"] = dict(reasons)
    stats["_lecture"] = (f"score et Elo du point de vue de {label_a}. "
                         f"Un Elo positif signifie que {label_a} est plus fort.")
    return stats


def run_ladder(snapshots, reference_ckpt, vocab, device, n_games,
               random_plies, seed, temperature):
    """Fait jouer chaque instantané contre le modèle de référence.

    Le résultat est un écart d'Elo par instantané, relatif au modèle final.
    L'origine est donc arbitraire — c'est une courbe de progression, pas une
    échelle absolue. Le seul Elo absolu qui vaudra quelque chose sera celui
    que Lichess attribuera au bot.
    """
    ref = ChessEngine(reference_ckpt, vocab, device, temperature)
    results = []

    for step, path in snapshots:
        eng = ChessEngine(path, vocab, device, temperature)
        rng = random.Random(seed + step)
        wins = draws = losses = 0
        reasons = Counter()
        t0 = time.perf_counter()

        for i in range(n_games):
            r, reason = play_game_model_vs_model(
                eng, ref, i % 2 == 0, random_plies, rng)
            reasons[reason] += 1
            if r == 1.0:
                wins += 1
            elif r == 0.5:
                draws += 1
            else:
                losses += 1

        stats = elo_with_error(wins, draws, losses)
        stats["step"] = step
        stats["tokens_vus"] = eng.tokens_seen
        stats["duree_s"] = round(time.perf_counter() - t0, 1)
        stats["terminaisons"] = dict(reasons)
        results.append(stats)
        e = stats["elo_diff"]
        print(f"  step {step:>6} ({eng.tokens_seen/1e6:6.1f} M tokens) : "
              f"score {stats['score']:.1%} "
              f"(+{wins} ={draws} -{losses}) | "
              f"Elo {e:+.0f}" if e is not None else
              f"  step {step:>6} : score {stats['score']:.1%}, Elo non borné",
              flush=True)
        del eng

    return results


# ---------------------------------------------------------------------------
# Un match complet contre un niveau donné
# ---------------------------------------------------------------------------

def run_match(model_engine, stockfish_path, skill_level, n_games, movetime_ms,
              random_plies, seed, verbose=True):
    rng = random.Random(seed)
    sf = chess.engine.SimpleEngine.popen_uci(stockfish_path)
    sf.configure({"Skill Level": skill_level, "Threads": 1, "Hash": 16})
    limit = chess.engine.Limit(time=movetime_ms / 1000.0)

    wins = draws = losses = 0
    reasons = Counter()
    t0 = time.perf_counter()

    try:
        for i in range(n_games):
            # On alterne les couleurs : jouer toujours blanc gonflerait le
            # score d'un avantage qui n'a rien à voir avec la force du modèle.
            model_is_white = (i % 2 == 0)
            result, reason = play_game(model_engine, sf, model_is_white, limit,
                                       random_plies, rng)
            reasons[reason] += 1
            if result == 1.0:
                wins += 1
            elif result == 0.5:
                draws += 1
            else:
                losses += 1

            if verbose and (i + 1) % 20 == 0:
                sc = (wins + 0.5 * draws) / (i + 1)
                print(f"    niveau {skill_level} : {i+1}/{n_games} parties | "
                      f"+{wins} ={draws} -{losses} | score {sc:.1%} | "
                      f"{time.perf_counter()-t0:.0f} s", flush=True)
    finally:
        sf.quit()

    stats = elo_with_error(wins, draws, losses)
    stats["skill_level"] = skill_level
    stats["movetime_ms"] = movetime_ms
    stats["random_plies"] = random_plies
    stats["duree_s"] = round(time.perf_counter() - t0, 1)
    stats["terminaisons"] = dict(reasons)
    return stats


# ---------------------------------------------------------------------------

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", required=True)
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--device", default="cuda:0")
    p.add_argument("--stockfish", default="stockfish")
    p.add_argument("--levels", type=int, nargs="+", default=[0, 1, 2, 3])
    p.add_argument("--games", type=int, default=200)
    p.add_argument("--movetime-ms", type=int, default=50)
    p.add_argument("--random-plies", type=int, default=4)
    p.add_argument("--temperature", type=float, default=0.0)
    p.add_argument("--seed", type=int, default=2026)
    p.add_argument("--out", default="logs/elo_report.json")
    p.add_argument("--ladder", action="store_true",
                   help="fait aussi jouer les instantanés contre --ckpt")
    p.add_argument("--ladder-dir", default="checkpoints")
    p.add_argument("--ladder-run", default="run1")
    p.add_argument("--ladder-games", type=int, default=60)
    p.add_argument("--skip-stockfish", action="store_true")
    p.add_argument("--vs", default=None,
                   help="checkpoint adverse pour un duel direct contre --ckpt")
    p.add_argument("--vs-games", type=int, default=300)
    p.add_argument("--label-a", default="A")
    p.add_argument("--label-b", default="B")
    args = p.parse_args()

    eng = ChessEngine(args.ckpt, args.vocab, args.device, args.temperature)
    print(f"[modèle] {args.ckpt} | step {eng.ckpt_step:,} | "
          f"{eng.tokens_seen:,} tokens vus")
    print(f"[adversaire] {args.stockfish}, {args.movetime_ms} ms par coup, "
          f"niveaux {args.levels}")
    print(f"[protocole] {args.games} parties par niveau, couleurs alternées, "
          f"{args.random_plies} demi-coups d'ouverture aléatoires\n")

    report = {
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "checkpoint": os.path.abspath(args.ckpt),
        "step": eng.ckpt_step,
        "tokens_vus": eng.tokens_seen,
        "protocole": {
            "parties_par_niveau": args.games,
            "movetime_ms": args.movetime_ms,
            "demi_coups_aleatoires": args.random_plies,
            "couleurs": "alternées, le modèle joue blanc une partie sur deux",
            "temperature_modele": args.temperature,
            "juge": ("python-chess, pas cutechess-cli — ce dernier n'est "
                     "empaqueté ni dans Homebrew ni dans apt sur cette machine "
                     "et sa compilation exige Qt6 et les droits administrateur"),
        },
        "matchs": [],
    }

    for level in ([] if args.skip_stockfish else args.levels):
        print(f"  == Stockfish Skill Level {level} ==", flush=True)
        stats = run_match(eng, args.stockfish, level, args.games,
                          args.movetime_ms, args.random_plies, args.seed)
        report["matchs"].append(stats)
        d = stats
        if d["elo_diff"] is None:
            print(f"    -> score {d['score']:.1%} "
                  f"(+{d['victoires']} ={d['nulles']} -{d['defaites']}) | "
                  f"écart d'Elo non borné (score extrême)\n")
        else:
            print(f"    -> score {d['score']:.1%} "
                  f"(+{d['victoires']} ={d['nulles']} -{d['defaites']}) | "
                  f"Elo {d['elo_diff']:+.0f} ± {d['marge_erreur']:.0f}\n")

    if args.vs:
        print(f"\n  == Duel direct : {args.label_a} contre {args.label_b}, "
              f"{args.vs_games} parties ==", flush=True)
        h2h = run_head_to_head(args.ckpt, args.vs, args.vocab, args.device,
                               args.vs_games, args.random_plies, args.seed,
                               args.temperature, args.label_a, args.label_b)
        report["duel_direct"] = h2h
        e, mg = h2h["elo_diff"], h2h["marge_erreur"]
        detail = (f"+{h2h['victoires']} ={h2h['nulles']} -{h2h['defaites']}")
        # La marge peut être indéfinie alors que l'Elo ne l'est pas : il suffit
        # que la borne haute de l'intervalle atteigne 100 % de score pour que
        # l'Elo correspondant diverge. Les deux valeurs se testent séparément.
        if e is None:
            print(f"    -> {args.label_a} marque {h2h['score']:.1%} ({detail}) | "
                  f"écart d'Elo non borné (score extrême)\n")
        elif mg is None:
            print(f"    -> {args.label_a} marque {h2h['score']:.1%} ({detail}) | "
                  f"Elo {e:+.0f}, marge non bornée (trop peu de parties)\n")
        else:
            print(f"    -> {args.label_a} marque {h2h['score']:.1%} ({detail}) | "
                  f"Elo {e:+.0f} ± {mg:.0f}\n")

    if args.ladder:
        import glob
        import re
        pat = os.path.join(args.ladder_dir, f"{args.ladder_run}_step*.pt")
        snaps = []
        for path in glob.glob(pat):
            m = re.search(r"_step(\d+)\.pt$", path)
            if m:
                snaps.append((int(m.group(1)), path))
        snaps.sort()
        print(f"\n  == Échelle : {len(snaps)} instantanés contre "
              f"{os.path.basename(args.ckpt)}, "
              f"{args.ladder_games} parties chacun ==", flush=True)
        ladder = run_ladder(snaps, args.ckpt, args.vocab, args.device,
                            args.ladder_games, args.random_plies, args.seed,
                            args.temperature)
        report["echelle_instantanes"] = {
            "_methode": ("chaque instantané joue contre le checkpoint de "
                         "référence passé en --ckpt, couleurs alternées, "
                         "ouvertures aléatoires. L'écart d'Elo est donc "
                         "relatif au modèle final : l'origine est arbitraire, "
                         "c'est la forme de la courbe qui a un sens."),
            "reference": os.path.abspath(args.ckpt),
            "parties_par_instantane": args.ladder_games,
            "points": ladder,
        }
        # Format directement consommable par plots.py
        report["courbe_elo_vs_tokens"] = [
            {"tokens_vus": p["tokens_vus"],
             "elo": p["elo_diff"],
             "marge_erreur": p["marge_erreur"]}
            for p in ladder
            if p["elo_diff"] is not None and p["marge_erreur"] is not None
        ]

    report["note_elo_absolu"] = (
        "Les chiffres ci-dessus sont des écarts d'Elo *relatifs* à Stockfish "
        "au niveau et à la cadence indiqués, mesurés sur cette machine. Les "
        "convertir en Elo absolu sur l'échelle Lichess demanderait un point "
        "d'ancrage que ce projet n'a pas vérifié lui-même. Le seul Elo absolu "
        "qui vaudra sera celui que Lichess attribuera au bot après ses "
        "premières parties classées, à la phase 6.")

    os.makedirs(os.path.dirname(args.out), exist_ok=True)
    with open(args.out, "w") as f:
        json.dump(report, f, indent=2, ensure_ascii=False)
    print(f"Rapport -> {args.out}")


if __name__ == "__main__":
    main()
qwen_match.py441 lines

Duel against a general-purpose LLM through its API, with the three protocols.

"""Notre modèle spécialisé contre un généraliste de 35 milliards de paramètres.

Le contraste est l'intérêt de l'exercice : d'un côté 142 M de paramètres
entraînés deux heures sur une seule tâche, de l'autre un modèle 250 fois plus
gros entraîné sur tout le web. Lequel joue le mieux aux échecs ?

### Ce qu'on mesure, et pourquoi deux choses séparément

**1. Le taux de coups légaux.** C'est la mesure directement comparable à notre
chiffre central : 97.86 % pour run1. On demande à Qwen un coup, on regarde s'il
est jouable, sans lui donner de seconde chance. Ce nombre-là est le plus
parlant de tous parce qu'il oppose les deux modèles sur exactement la même
question.

**2. Le résultat des parties.** Là, on accorde des tentatives (`--retries`).
Sans cela, presque toutes les parties se termineraient au dixième coup sur une
illégalité, et on ne mesurerait rien de la qualité du jeu. Les deux chiffres
répondent à deux questions différentes et sont publiés séparément.

### Les précautions de protocole

Un généraliste doit être interrogé correctement, sinon on mesure sa capacité à
suivre une consigne plutôt que sa force aux échecs. Trois précautions :

- on lui donne la partie en UCI **et** la liste des coups légaux n'est PAS
  fournie — la fournir reviendrait à lui offrir le masque que notre modèle
  n'a pas en génération libre ;
- on accepte sa réponse en UCI ou en notation courante, et on extrait le coup
  même s'il l'enrobe de phrases ou de raisonnement ;
- la température est celle de son service (0.6), pas 0 : c'est le régime dans
  lequel ce modèle est conçu pour fonctionner.

Usage :
    python qwen_match.py --ckpt checkpoints/run2_best.pt --games 100
"""

import argparse
import json
import os
import random
import re
import time
import urllib.error
import urllib.request
from collections import Counter

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import chess

from elo_match import elo_with_error
from engine import ChessEngine

SYSTEM = (
    "Tu es un moteur d'échecs. On te donne les coups déjà joués d'une partie, "
    "en notation UCI (case de départ puis case d'arrivée, par exemple e2e4). "
    "Tu réponds UNIQUEMENT par le meilleur coup suivant, en notation UCI, "
    "sans aucun autre mot, sans ponctuation, sans explication."
)

UCI_RE = re.compile(r"\b([a-h][1-8][a-h][1-8][qrbn]?)\b")
THINK_RE = re.compile(r"<think>.*?</think>", re.S)


def construire_message(historique, refuses, seuil_hasard):
    """Compose la demande, en escaladant si les tentatives précédentes ont échoué.

    Trois régimes successifs, et le passage de l'un à l'autre est le coeur de
    la mesure :

    1. Première tentative : on demande simplement le meilleur coup. C'est cette
       réponse-là, et elle seule, qui compte pour le taux de légalité comparé
       à notre modèle.
    2. Tentatives suivantes : on lui dit que ses propositions précédentes
       étaient illégales et on les liste. Il dispose donc d'une information
       que notre modèle n'a jamais — c'est délibérément généreux.
    3. Au-delà du seuil : on abandonne l'exigence de qualité et on lui demande
       n'importe quel coup légal, même au hasard. S'il échoue encore là, ce
       n'est plus une question de force de jeu.
    """
    if historique:
        base = ("Coups joués : " + " ".join(historique)
                + "\n\nQuel est le meilleur coup suivant ?")
    else:
        base = "La partie commence. Quel est ton premier coup ?"

    if not refuses:
        return base

    liste = ", ".join(refuses)
    if len(refuses) < seuil_hasard:
        return (base + f"\n\nATTENTION : tu as déjà proposé {liste}. "
                f"Ces coups sont ILLÉGAUX dans cette position. "
                f"Propose un coup différent et réellement jouable.")
    return (base + f"\n\nTu as déjà proposé {len(refuses)} coups illégaux : "
            f"{liste}.\n"
            f"Oublie la qualité du coup. Donne n'importe quel coup LÉGAL, "
            f"même complètement au hasard. Le seul critère est qu'il soit "
            f"jouable dans cette position.")


def interroger(url, modele, contenu, temperature, timeout, max_tokens,
               sans_raisonnement=True):
    """Envoie la demande au serveur et renvoie le texte brut de la réponse."""

    payload = {
        "model": modele,
        "messages": [{"role": "system", "content": SYSTEM},
                     {"role": "user", "content": contenu}],
        "temperature": temperature,
        "max_tokens": max_tokens,
    }
    # Qwen3 est un modele de raisonnement : par defaut il produit un bloc de
    # reflexion de ~530 tokens avant de repondre, soit 7.7 s par coup, donc
    # neuf heures pour cent parties. enable_thinking=false lui fait donner la
    # meme reponse en 5 tokens et 0.4 s, soit dix-neuf fois plus vite.
    #
    # Ce n'est pas un bridage deloyal : la reponse mesuree est identique sur
    # les positions testees. Et le protocole reste plus genereux que celui de
    # notre modele, qui n'a droit qu'a un seul passage avant sans reflexion.
    if sans_raisonnement:
        payload["chat_template_kwargs"] = {"enable_thinking": False}
    corps = json.dumps(payload).encode()

    req = urllib.request.Request(
        url, data=corps, headers={"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=timeout) as r:
        d = json.loads(r.read())
    return d["choices"][0]["message"]["content"]


def extraire_coup(texte, board):
    """Retrouve un coup dans la réponse, en UCI ou en notation courante.

    Les modèles de raisonnement enveloppent souvent leur réponse dans un bloc
    <think>. On le retire d'abord : le coup qu'on veut est celui de la
    conclusion, pas ceux évoqués pendant la réflexion.
    """
    texte = THINK_RE.sub(" ", texte).strip()

    # On parcourt les candidats UCI en partant de la FIN : si le modèle a
    # énuméré plusieurs coups avant de conclure, sa conclusion est en dernier.
    for m in reversed(UCI_RE.findall(texte)):
        try:
            coup = chess.Move.from_uci(m)
        except ValueError:
            continue
        if coup in board.legal_moves:
            return coup, m, "uci_legal"
        return coup, m, "uci_illegal"

    # Repli sur la notation courante, au cas où il l'utilise malgré la consigne.
    for mot in reversed(texte.replace(",", " ").replace(".", " ").split()):
        mot = mot.strip("*`\"'()[]")
        try:
            coup = board.parse_san(mot)
            return coup, mot, "san_legal"
        except ValueError:
            continue

    return None, texte[:60], "illisible"


def classer_illegal(coup, board):
    """Pourquoi ce coup est-il impossible ? Même typologie que pour notre modèle."""
    p = board.piece_at(coup.from_square)
    if p is None:
        return "case_depart_vide"
    if p.color != board.turn:
        return "piece_adverse"
    return "deplacement_illegal"


def jouer_partie(eng, url, modele, notre_couleur_blanche, rng, args, stats):
    """Une partie. Renvoie le résultat du point de vue de NOTRE modèle."""
    board = chess.Board()
    historique = []

    for _ in range(args.random_plies):
        if board.is_game_over():
            break
        c = rng.choice(list(board.legal_moves))
        board.push(c)
        historique.append(c.uci())

    while not board.is_game_over(claim_draw=True) and len(historique) < args.max_plies:
        a_nous = (board.turn == chess.WHITE) == notre_couleur_blanche

        if a_nous:
            eng.set_position(board, historique)
            coup = eng.best_move(board)
        else:
            coup = None
            refuses = []
            for essai in range(args.max_attempts):
                contenu = construire_message(historique, refuses,
                                             args.seuil_hasard)
                try:
                    texte = interroger(url, modele, contenu, args.temperature,
                                       args.timeout, args.max_tokens,
                                       not args.avec_raisonnement)
                except (urllib.error.URLError, OSError, KeyError):
                    stats["erreurs_reseau"] += 1
                    time.sleep(2)
                    continue
                c, brut, genre = extraire_coup(texte, board)

                # Le taux de légalité ne compte QUE la première tentative.
                # Accorder des essais puis mesurer le meilleur fausserait la
                # comparaison avec notre modèle, qui n'en a qu'un.
                if essai == 0:
                    stats["premiere_tentative"] += 1
                    stats["p1_" + genre] += 1
                    if c is not None and c not in board.legal_moves:
                        stats["p1_type_" + classer_illegal(c, board)] += 1

                if c is not None and c in board.legal_moves:
                    coup = c
                    stats["coups_obtenus"] += 1
                    stats["total_tentatives"] += essai + 1
                    stats[f"essais_{min(essai + 1, 20)}"] += 1
                    if essai + 1 >= args.seuil_hasard:
                        stats["obtenus_apres_seuil_hasard"] += 1
                    break
                if c is not None:
                    refuses.append(brut)
                    stats["propositions_illegales"] += 1
                    if essai + 1 >= args.seuil_hasard:
                        stats["illegales_apres_seuil_hasard"] += 1
                else:
                    refuses.append("(illisible)")

            if coup is None:
                # Même avec toutes les tentatives et la consigne de jouer au
                # hasard, aucun coup jouable. On joue un coup légal à sa place
                # plutôt que d'arrêter la partie : le but est de mesurer son
                # jeu, pas de compter les forfaits.
                stats["echecs_totaux"] += 1
                stats["total_tentatives"] += args.max_attempts
                if args.forfait_si_echec:
                    return (1.0, "adversaire_illegal")
                coup = rng.choice(list(board.legal_moves))
                stats["coups_joues_a_sa_place"] += 1

        board.push(coup)
        historique.append(coup.uci())

    if len(historique) >= args.max_plies:
        return 0.5, "limite_de_coups"
    issue = board.outcome(claim_draw=True)
    if issue is None or issue.winner is None:
        return 0.5, ("indetermine" if issue is None
                     else issue.termination.name.lower())
    gagne = (issue.winner == chess.WHITE) == notre_couleur_blanche
    return (1.0 if gagne else 0.0), issue.termination.name.lower()


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", default="checkpoints/run2_best.pt")
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--device", default="cuda:0")
    p.add_argument("--url", default="http://localhost:8080/v1/chat/completions")
    p.add_argument("--modele", default="qwen36-35b")
    p.add_argument("--label", default="Qwen3.6-35B-A3B")
    p.add_argument("--games", type=int, default=100)
    p.add_argument("--max-attempts", type=int, default=30,
                   help="tentatives accordées à l'adversaire pour trouver un coup légal")
    p.add_argument("--seuil-hasard", type=int, default=10,
                   help="au-delà, on lui demande n'importe quel coup légal")
    p.add_argument("--forfait-si-echec", action="store_true",
                   help="perdre la partie après max-attempts, au lieu de jouer "
                        "un coup légal à sa place")
    p.add_argument("--strict", action="store_true",
                   help="RÈGLE RÉELLE : le premier coup illégal fait perdre la "
                        "partie, comme en tournoi. Équivaut à --max-attempts 1 "
                        "--forfait-si-echec.")
    p.add_argument("--temperature", type=float, default=0.6)
    p.add_argument("--random-plies", type=int, default=2)
    p.add_argument("--max-plies", type=int, default=200)
    p.add_argument("--max-tokens", type=int, default=256)
    p.add_argument("--avec-raisonnement", action="store_true",
                   help="laisse Qwen produire son bloc de reflexion, 19x plus lent")
    p.add_argument("--timeout", type=float, default=180)
    p.add_argument("--seed", type=int, default=99)
    p.add_argument("--out", default="logs/duel_qwen.json")
    args = p.parse_args()
    if args.strict:
        # En tournoi, un moteur qui répond un coup illégal perd immédiatement.
        # C'est la règle FIDE et celle de tout arbitre entre moteurs. Notre
        # propre moteur ne peut jamais s'y exposer : engine.py masque les coups
        # illégaux avant de choisir. Ce mode applique donc à l'adversaire
        # exactement la contrainte que subirait n'importe quel participant.
        args.max_attempts = 1
        args.forfait_si_echec = True

    # Vérification préalable : sans le serveur, autant s'arrêter tout de suite.
    sante = args.url.rsplit("/v1/", 1)[0] + "/health"
    try:
        with urllib.request.urlopen(sante, timeout=10) as r:
            r.read()
    except Exception as e:
        raise SystemExit(
            f"Serveur injoignable sur {sante} ({e}).\n"
            f"Le démarrer avec : sudo systemctl start llama.service")

    eng = ChessEngine(args.ckpt, args.vocab, args.device, temperature=0.0)
    print(f"[notre modèle] {os.path.basename(args.ckpt)} — step {eng.ckpt_step:,}, "
          f"{eng.tokens_seen/1e6:.0f} M tokens vus")
    print(f"[adversaire]   {args.label}, température {args.temperature}, "
          f"jusqu'à {args.max_attempts} tentatives, consigne de jouer "
          f"au hasard à partir de {args.seuil_hasard}")
    print(f"[protocole]    {args.games} parties, couleurs alternées, "
          f"{args.random_plies} demi-coups d'ouverture aléatoires\n")

    rng = random.Random(args.seed)
    stats = Counter()
    v = n = d_ = 0
    fins = Counter()
    t0 = time.perf_counter()

    for i in range(args.games):
        r, raison = jouer_partie(eng, args.url, args.modele, i % 2 == 0,
                                 rng, args, stats)
        fins[raison] += 1
        if r == 1.0:
            v += 1
        elif r == 0.5:
            n += 1
        else:
            d_ += 1
        if (i + 1) % 5 == 0:
            n1 = max(1, stats["premiere_tentative"])
            leg = (stats["p1_uci_legal"] + stats["p1_san_legal"]) / n1
            moy = stats["total_tentatives"] / max(1, stats["coups_obtenus"]
                                                  + stats["echecs_totaux"])
            print(f"  {i+1}/{args.games} | +{v} ={n} -{d_} | "
                  f"légaux au 1er essai : {leg:.1%} | "
                  f"{moy:.1f} tentatives/coup | "
                  f"{time.perf_counter()-t0:.0f} s", flush=True)

    from evaluate import wilson_interval
    n1 = stats["premiere_tentative"]
    legaux1 = stats["p1_uci_legal"] + stats["p1_san_legal"]
    taux, lo, hi = (wilson_interval(legaux1, n1) if n1 else (None, None, None))

    obtenus = stats["coups_obtenus"]
    moy = (stats["total_tentatives"] / max(1, obtenus + stats["echecs_totaux"]))
    distribution = {str(k): stats[f"essais_{k}"] for k in range(1, 21)
                    if stats[f"essais_{k}"]}

    rapport = {
        "date_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "notre_modele": {"checkpoint": os.path.abspath(args.ckpt),
                         "step": eng.ckpt_step, "tokens_vus": eng.tokens_seen},
        "adversaire": {"nom": args.label, "temperature": args.temperature,
                       "max_tentatives_par_coup": args.max_attempts,
                       "seuil_consigne_hasard": args.seuil_hasard,
                       "raisonnement_active": args.avec_raisonnement,
                       "note_raisonnement": (
                           "le bloc de reflexion de Qwen3 est desactive : il "
                           "coute 7.7 s par coup contre 0.4 s, pour une reponse "
                           "identique sur les positions testees")},
        "protocole": {
            "parties": args.games,
            "couleurs": "alternées",
            "demi_coups_aleatoires": args.random_plies,
            "note_legalite": ("le taux de légalité ne compte QUE la première "
                              "tentative de chaque coup ; accorder des essais "
                              "puis mesurer le meilleur fausserait la comparaison "
                              "avec notre modèle, qui n'en a qu'un"),
            "note_masque": ("la liste des coups légaux n'est PAS fournie à "
                            "l'adversaire — la fournir reviendrait à lui offrir "
                            "le masque que notre modèle n'a pas non plus en "
                            "génération libre")},
        "legalite_premiere_tentative": {
            "_role": ("LE chiffre comparable à notre modèle : même question, "
                      "aucune liste de coups fournie, aucune seconde chance"),
            "coups_demandes": n1,
            "legaux": legaux1,
            "taux": None if taux is None else round(taux, 5),
            "ic95_bas": None if lo is None else round(lo, 5),
            "ic95_haut": None if hi is None else round(hi, 5),
            "format_de_reponse": {k[3:]: stats[k] for k in
                                  ("p1_uci_legal", "p1_uci_illegal",
                                   "p1_san_legal", "p1_illisible")},
            "typologie_des_coups_impossibles": {
                k[9:]: stats[k] for k in
                ("p1_type_case_depart_vide", "p1_type_piece_adverse",
                 "p1_type_deplacement_illegal") if stats[k]}},

        "insistance_necessaire": {
            "_role": ("combien de relances faut-il pour obtenir un coup jouable ? "
                      "À chaque échec on lui dit que son coup était illégal et on "
                      "lui liste ses propositions refusées — une information que "
                      "notre modèle n'a jamais"),
            "coups_finalement_obtenus": obtenus,
            "tentatives_totales": stats["total_tentatives"],
            "tentatives_moyennes_par_coup": round(moy, 2),
            "distribution_tentatives": distribution,
            "propositions_illegales_totales": stats["propositions_illegales"],
            "seuil_hasard": args.seuil_hasard,
            "obtenus_apres_seuil_hasard": stats["obtenus_apres_seuil_hasard"],
            "illegales_apres_seuil_hasard": stats["illegales_apres_seuil_hasard"],
            "echecs_malgre_tout": stats["echecs_totaux"],
            "coups_joues_a_sa_place": stats["coups_joues_a_sa_place"],
            "max_attempts": args.max_attempts,
            "erreurs_reseau": stats["erreurs_reseau"]},
        "resultat_parties": elo_with_error(v, n, d_),
        "terminaisons": dict(fins),
        "duree_s": round(time.perf_counter() - t0, 1),
    }
    rapport["resultat_parties"]["_lecture"] = (
        "score et Elo du point de vue de NOTRE modèle ; un Elo positif signifie "
        "que notre modèle est plus fort")

    os.makedirs(os.path.dirname(args.out), exist_ok=True)
    with open(args.out, "w") as f:
        json.dump(rapport, f, indent=2, ensure_ascii=False)

    print(f"\n=== {args.games} parties en {rapport['duree_s']/60:.0f} min ===")
    print(f"notre modèle : +{v} ={n} -{d_}  "
          f"soit {rapport['resultat_parties']['score']:.1%}")
    print(f"fins de partie : {dict(fins)}")
    if taux is not None:
        print(f"\n{args.label}, PREMIÈRE tentative : {taux:.2%} de coups légaux "
              f"[IC95 {lo:.2%} - {hi:.2%}] sur {n1:,} coups")
    print(f"  tentatives moyennes pour obtenir un coup jouable : {moy:.2f}")
    print(f"  propositions illégales au total : {stats['propositions_illegales']:,}")
    print(f"  distribution : {distribution}")
    if stats["illegales_apres_seuil_hasard"]:
        print(f"  encore {stats['illegales_apres_seuil_hasard']} illégales APRÈS "
              f"qu'on lui ait demandé un coup au hasard")
    if stats["echecs_totaux"]:
        print(f"  {stats['echecs_totaux']} coups jamais trouvés en "
              f"{args.max_attempts} tentatives")
    print(f"\nrapport -> {args.out}")


if __name__ == "__main__":
    main()
play.py168 lines

Play against the model from the keyboard, with its opinion displayed.

"""Jouer contre le modèle, dans le terminal.

    python play.py                                  # partie contre run1, tu joues blanc
    python play.py --ckpt checkpoints/run2_best.pt  # contre run2
    python play.py --noir                           # tu joues noir
    python play.py --temperature 0.4                # un adversaire moins prévisible

Les coups s'entrent indifféremment en notation courante (`Cf3`, `e4`, `O-O`) ou
en UCI (`g1f3`, `e2e4`, `e1g1`). Commandes disponibles pendant la partie :

    ?          la liste des coups légaux
    top        ce que le modèle pense de la position actuelle
    annuler    revenir en arrière d'un coup complet
    abandon    terminer la partie
"""

import argparse
import os

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import chess

from engine import ChessEngine

PIECES = {
    "P": "♙", "N": "♘", "B": "♗", "R": "♖", "Q": "♕", "K": "♔",
    "p": "♟", "n": "♞", "b": "♝", "r": "♜", "q": "♛", "k": "♚",
}


def afficher(board: chess.Board, vue_blanc: bool = True):
    """Dessine l'échiquier, orienté du côté du joueur humain."""
    rangs = range(7, -1, -1) if vue_blanc else range(8)
    cols = range(8) if vue_blanc else range(7, -1, -1)
    print()
    for r in rangs:
        ligne = f"  {r + 1} "
        for c in cols:
            p = board.piece_at(chess.square(c, r))
            ligne += (PIECES[p.symbol()] if p else "·") + " "
        print(ligne)
    lettres = "abcdefgh" if vue_blanc else "hgfedcba"
    print("    " + " ".join(lettres))
    if board.is_check():
        print("\n  ÉCHEC")
    print()


def afficher_analyse(eng: ChessEngine, board: chess.Board, n=6):
    """Montre les coups que le modèle juge les plus probables.

    C'est la fenêtre la plus directe sur ce que le modèle a appris : les
    probabilités affichées sont celles qu'il attribue *après* masquage des
    coups illégaux, donc renormalisées sur les seuls coups jouables.
    """
    scores = eng.move_scores(board)
    top = sorted(scores.items(), key=lambda kv: -kv[1])[:n]
    print("  ce que le modèle envisage :")
    for uci, p in top:
        san = board.san(chess.Move.from_uci(uci))
        barre = "█" * max(1, round(p * 30))
        print(f"    {san:<8} {uci}  {p:6.1%} {barre}")
    print()


def lire_coup(board: chess.Board) -> chess.Move | str:
    """Lit un coup au clavier, en notation courante ou en UCI."""
    while True:
        txt = input("  ton coup > ").strip()
        if not txt:
            continue
        bas = txt.lower()
        if bas in ("?", "aide", "coups"):
            print("  " + " ".join(sorted(board.san(m) for m in board.legal_moves)))
            continue
        if bas in ("top", "analyse"):
            return "analyse"
        if bas in ("annuler", "undo"):
            return "annuler"
        if bas in ("abandon", "quit", "q"):
            return "abandon"
        # On tente d'abord la notation courante, puis l'UCI.
        try:
            return board.parse_san(txt)
        except ValueError:
            pass
        try:
            m = chess.Move.from_uci(bas)
            if m in board.legal_moves:
                return m
            print(f"  {txt} n'est pas légal ici. Tape ? pour la liste.")
        except ValueError:
            print(f"  {txt} n'est ni une notation courante ni un coup UCI.")


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", default="checkpoints/run1_best.pt")
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--device", default="cuda:0")
    p.add_argument("--temperature", type=float, default=0.0,
                   help="0 = le modèle joue toujours son meilleur coup")
    p.add_argument("--top-k", type=int, default=None)
    p.add_argument("--noir", action="store_true", help="tu joues les noirs")
    p.add_argument("--analyse", action="store_true",
                   help="montre l'avis du modèle à chaque coup")
    args = p.parse_args()

    eng = ChessEngine(args.ckpt, args.vocab, args.device,
                      args.temperature, args.top_k)
    humain_blanc = not args.noir

    print(f"\n  Modèle : {os.path.basename(args.ckpt)} "
          f"(step {eng.ckpt_step:,}, {eng.tokens_seen/1e6:.0f} M tokens vus)")
    print(f"  Tu joues les {'noirs' if args.noir else 'blancs'}. "
          f"Tape ? pour l'aide.\n")

    board = chess.Board()
    histoire: list[str] = []

    while not board.is_game_over(claim_draw=True):
        tour_humain = (board.turn == chess.WHITE) == humain_blanc
        afficher(board, humain_blanc)

        if tour_humain:
            if args.analyse:
                eng.set_position(board, histoire)
                afficher_analyse(eng, board)
            coup = lire_coup(board)
            if coup == "abandon":
                print("  Tu abandonnes.")
                return
            if coup == "analyse":
                eng.set_position(board, histoire)
                afficher_analyse(eng, board, n=10)
                continue
            if coup == "annuler":
                # Deux demi-coups : le sien et celui du modèle.
                for _ in range(2):
                    if board.move_stack:
                        board.pop()
                        histoire.pop()
                continue
            board.push(coup)
            histoire.append(coup.uci())
        else:
            eng.set_position(board, histoire)
            coup = eng.best_move(board)
            san = board.san(coup)
            board.push(coup)
            histoire.append(coup.uci())
            print(f"  Le modèle joue : {san}  ({coup.uci()})")

    afficher(board, humain_blanc)
    issue = board.outcome(claim_draw=True)
    if issue.winner is None:
        print(f"  Partie nulle — {issue.termination.name.lower()}")
    else:
        gagnant = "toi" if (issue.winner == chess.WHITE) == humain_blanc \
                  else "le modèle"
        print(f"  {gagnant.capitalize()} gagne — "
              f"{issue.termination.name.lower()}")
    print(f"\n  {len(histoire)} demi-coups : {' '.join(histoire)}\n")


if __name__ == "__main__":
    main()
convert_to_hf.py199 lines

Export to the Hugging Face format, verified numerically before writing.

"""Convertit un checkpoint du projet au format HuggingFace `LlamaForCausalLM`.

Pourquoi Llama, alors qu'on n'a rien copié de Llama ? Parce qu'en choisissant
séparément RMSNorm, RoPE, SwiGLU et le pre-norm — quatre choix qui se
justifient chacun indépendamment — on a reconstruit exactement l'architecture
de Llama. La correspondance n'est donc pas une approximation : c'est le même
calcul, écrit autrement.

Ce que ça débloque, et c'est l'essentiel :

* le modèle se charge avec `transformers` sans une ligne de notre code ;
* il se publie sur HuggingFace comme n'importe quel modèle Llama ;
* et surtout, `convert_hf_to_gguf.py` de llama.cpp sait le convertir en GGUF,
  donc le modèle devient utilisable par `llama-server`, `llama-cli`, Ollama,
  LM Studio et tout ce qui parle GGUF.

Le script vérifie systématiquement que les deux implémentations produisent les
mêmes logits, à la précision flottante près. Une conversion qui « a l'air de
marcher » mais décale d'un demi-pourcent produirait un modèle silencieusement
dégradé — le genre d'erreur qu'on ne remarque qu'en comparant les Elo.

Usage :
    python convert_to_hf.py --ckpt checkpoints/run1_best.pt --out hf/chess-51m
"""

import argparse
import json
import os
import shutil

os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")

import torch

from model import ChessGPT, ModelConfig


def build_state_dict(src: dict, cfg: ModelConfig) -> dict:
    """Renomme et redécoupe les poids vers la nomenclature de Llama.

    Une seule opération est non triviale : notre projection q/k/v est une
    matrice unique de taille 3*d x d, fusionnée pour n'avoir qu'un gros produit
    matriciel au lieu de trois petits. Llama garde les trois séparées, il faut
    donc la redécouper — dans le bon ordre, qui est celui du `split` de notre
    passe avant.
    """
    d = cfg.n_embd
    out = {}
    out["model.embed_tokens.weight"] = src["tok_emb.weight"]

    for i in range(cfg.n_layer):
        p = f"blocks.{i}."
        q = f"model.layers.{i}."
        qkv = src[p + "attn.qkv.weight"]
        wq, wk, wv = qkv.split(d, dim=0)
        out[q + "self_attn.q_proj.weight"] = wq
        out[q + "self_attn.k_proj.weight"] = wk
        out[q + "self_attn.v_proj.weight"] = wv
        out[q + "self_attn.o_proj.weight"] = src[p + "attn.proj.weight"]
        out[q + "mlp.gate_proj.weight"] = src[p + "mlp.gate.weight"]
        out[q + "mlp.up_proj.weight"] = src[p + "mlp.up.weight"]
        out[q + "mlp.down_proj.weight"] = src[p + "mlp.down.weight"]
        out[q + "input_layernorm.weight"] = src[p + "norm_attn.weight"]
        out[q + "post_attention_layernorm.weight"] = src[p + "norm_mlp.weight"]

    out["model.norm.weight"] = src["norm_final.weight"]
    if not cfg.tie_embeddings:
        out["lm_head.weight"] = src["lm_head.weight"]
    return out


def make_tokenizer_files(vocab_path: str, out_dir: str):
    """Écrit un tokenizer au niveau du mot, où un « mot » est un coup UCI.

    C'est le point le plus inhabituel de cette conversion. Les modèles de
    langage courants découpent le texte en fragments de mots ; ici le
    vocabulaire est fermé — 1971 entrées, une par coup possible — et la
    séparation se fait sur les espaces. Le format `WordLevel` de la
    bibliothèque `tokenizers` correspond exactement à ce besoin.

    Conséquence pratique : le modèle se pilote en écrivant les coups séparés
    par des espaces, par exemple « e2e4 e7e5 g1f3 », et il complète la suite.
    """
    from tokenizers import Tokenizer, models, pre_tokenizers, decoders

    with open(vocab_path) as f:
        v = json.load(f)
    stoi = v["stoi"]

    tok = Tokenizer(models.WordLevel(vocab=stoi, unk_token="<pad>"))
    tok.pre_tokenizer = pre_tokenizers.WhitespaceSplit()
    # Le décodeur doit rejoindre les coups par des espaces. `Fuse` les
    # concatène sans séparateur, ce qui produit une chaîne du type
    # « e2e4e7e5g1f3 » indécodable : les coups UCI faisant quatre ou cinq
    # caractères, on ne peut plus les redécouper de façon fiable.
    # `WordPiece` avec un préfixe absent de notre vocabulaire fait exactement
    # ce qu'il faut : il joint par des espaces sans rien retirer.
    tok.decoder = decoders.WordPiece(prefix=chr(32), cleanup=False)
    tok.save(os.path.join(out_dir, "tokenizer.json"))

    with open(os.path.join(out_dir, "tokenizer_config.json"), "w") as f:
        json.dump({
            "tokenizer_class": "PreTrainedTokenizerFast",
            "model_max_length": 256,
            "bos_token": "<bos>",
            "eos_token": "<eos>",
            "pad_token": "<pad>",
            "unk_token": "<pad>",
            "clean_up_tokenization_spaces": False,
        }, f, indent=2)

    with open(os.path.join(out_dir, "special_tokens_map.json"), "w") as f:
        json.dump({"bos_token": "<bos>", "eos_token": "<eos>",
                   "pad_token": "<pad>", "unk_token": "<pad>"}, f, indent=2)


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--ckpt", required=True)
    p.add_argument("--vocab", default="data/vocab.json")
    p.add_argument("--out", required=True)
    p.add_argument("--tolerance", type=float, default=1e-4,
                   help="écart maximal toléré sur les logits")
    args = p.parse_args()

    from transformers import LlamaConfig, LlamaForCausalLM

    ck = torch.load(args.ckpt, map_location="cpu", weights_only=False)
    cfg = ModelConfig(**ck["model_config"])
    print(f"source : {args.ckpt}")
    print(f"  {cfg.n_layer} couches, d={cfg.n_embd}, {cfg.n_head} têtes, "
          f"MLP {cfg.mlp_hidden}, vocabulaire {cfg.vocab_size}")

    hf_cfg = LlamaConfig(
        vocab_size=cfg.vocab_size,
        hidden_size=cfg.n_embd,
        intermediate_size=cfg.mlp_hidden,
        num_hidden_layers=cfg.n_layer,
        num_attention_heads=cfg.n_head,
        num_key_value_heads=cfg.n_head,     # pas de GQA : autant de clés que de requêtes
        max_position_embeddings=cfg.block_size,
        rms_norm_eps=1e-6,                  # identique à notre RMSNorm
        rope_theta=cfg.rope_theta,
        tie_word_embeddings=cfg.tie_embeddings,
        attention_bias=False,
        mlp_bias=False,
        hidden_act="silu",
        bos_token_id=1, eos_token_id=2, pad_token_id=0,
    )

    hf = LlamaForCausalLM(hf_cfg)
    missing, unexpected = hf.load_state_dict(
        build_state_dict(ck["model"], cfg), strict=False)
    # Avec les embeddings liés, lm_head.weight n'est pas dans le dictionnaire :
    # transformers le relie lui-même à embed_tokens. C'est la seule absence
    # acceptable.
    assert not unexpected, f"poids inattendus : {unexpected}"
    assert all("lm_head" in m for m in missing), f"poids manquants : {missing}"
    hf.tie_weights()

    # --- Vérification numérique ---------------------------------------------
    ours = ChessGPT(cfg)
    ours.load_state_dict(ck["model"])
    ours.eval()
    hf.eval()

    torch.manual_seed(0)
    max_err = 0.0
    with torch.no_grad():
        for T in (1, 7, 64, cfg.block_size):
            x = torch.randint(0, cfg.vocab_size, (3, T))
            a, _ = ours(x)
            b = hf(x).logits
            err = (a - b).abs().max().item()
            max_err = max(max_err, err)
            print(f"  longueur {T:>3} : écart max sur les logits = {err:.2e}")

    if max_err > args.tolerance:
        raise SystemExit(
            f"\nECHEC : écart de {max_err:.2e}, au-dessus de la tolérance "
            f"{args.tolerance:.0e}. La conversion est incorrecte, ne pas publier.")
    print(f"\n[OK] écart maximal {max_err:.2e} — les deux implémentations "
          f"calculent la même chose.")

    # --- Écriture -----------------------------------------------------------
    os.makedirs(args.out, exist_ok=True)
    hf.save_pretrained(args.out, safe_serialization=True)
    make_tokenizer_files(args.vocab, args.out)
    shutil.copy(args.vocab, os.path.join(args.out, "vocab_uci.json"))

    n = sum(p.numel() for p in hf.parameters())
    print(f"\nÉcrit dans {args.out}/ ({n/1e6:.1f} M paramètres)")
    for f in sorted(os.listdir(args.out)):
        sz = os.path.getsize(os.path.join(args.out, f))
        print(f"  {f:<32} {sz/1e6:8.2f} Mo")


if __name__ == "__main__":
    main()
plots.py624 lines

The training progression curves.

"""Génération des graphiques PNG du projet.

Contrainte de lecture : ces images seront regardées sur mobile, dans un article
LinkedIn. D'où des polices grandes, peu de séries par graphique, des étiquettes
posées directement sur les courbes plutôt qu'une légende à décoder, et un
quadrillage discret qui ne concurrence pas les données.

Couleurs : les deux premiers slots de la palette catégorielle de référence
(bleu, orange), vérifiés avec le validateur de palette. On ne dépasse jamais
deux séries par graphique — au-delà, l'écart perceptuel entre courbes n'est
plus garanti en vision des couleurs déficiente.

Le texte porte toujours les jetons d'encre, jamais la couleur de la série :
c'est la courbe adjacente à l'étiquette qui porte l'identité, pas l'étiquette
elle-même. Une étiquette colorée serait illisible pour qui distingue mal les
teintes, alors que « validation 0.039 » posé au bout de la courbe orange reste
compréhensible sans percevoir l'orange.

Usage :
    python plots.py --run-name run1
"""

import argparse
import csv
import glob
import json
import os

import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib.ticker import FuncFormatter

# --- Jetons de style -------------------------------------------------------

SURFACE = "#fcfcfb"
INK_PRIMARY = "#0b0b0b"
INK_SECONDARY = "#52514e"
INK_MUTED = "#8a8984"
GRID = "#e6e5e1"

# Deux slots seulement : bleu et orange. Vérifié avec le validateur de palette
# (mode clair, tous appariements) : écart perceptuel 33.6 en vision normale et
# 24.7 en vision déficiente, contraste supérieur à 3:1 sur le fond pour les
# deux. Un troisième slot (aqua #1baf7a) avait été envisagé mais son contraste
# tombe à 2.74 sur ce fond clair, sous le seuil — écarté plutôt que rattrapé.
SERIES = ["#2a78d6", "#eb6834"]   # bleu, orange

plt.rcParams.update({
    "figure.facecolor": SURFACE,
    "axes.facecolor": SURFACE,
    "savefig.facecolor": SURFACE,
    "font.size": 13,
    "axes.titlesize": 17,
    "axes.labelsize": 14,
    "xtick.labelsize": 12,
    "ytick.labelsize": 12,
    "legend.fontsize": 13,
    "axes.edgecolor": GRID,
    "axes.labelcolor": INK_SECONDARY,
    "text.color": INK_PRIMARY,
    "xtick.color": INK_SECONDARY,
    "ytick.color": INK_SECONDARY,
    "grid.color": GRID,
    "grid.linewidth": 1.0,
    "lines.linewidth": 2.0,
    "figure.dpi": 160,
})


def style_axes(ax, title, xlabel, ylabel, subtitle=None):
    # Le titre est posé au-dessus du sous-titre, tous deux en coordonnées
    # d'axes : c'est le seul moyen fiable d'éviter que les jambages du titre
    # ne mordent sur la ligne du dessous quelle que soit la taille de figure.
    if subtitle:
        ax.text(0.0, 1.13, title, transform=ax.transAxes, color=INK_PRIMARY,
                fontsize=17, fontweight="bold", va="bottom", ha="left")
        ax.text(0.0, 1.03, subtitle, transform=ax.transAxes,
                color=INK_SECONDARY, fontsize=12, va="bottom", ha="left")
    else:
        ax.text(0.0, 1.03, title, transform=ax.transAxes, color=INK_PRIMARY,
                fontsize=17, fontweight="bold", va="bottom", ha="left")
    ax.set_xlabel(xlabel)
    ax.set_ylabel(ylabel)
    ax.grid(True, axis="y", linewidth=1.0, alpha=0.9)
    ax.set_axisbelow(True)
    # Quadrillage discret : on retire les bordures qui n'apportent rien.
    for side in ("top", "right"):
        ax.spines[side].set_visible(False)
    for side in ("left", "bottom"):
        ax.spines[side].set_color(GRID)


def millions(x, _):
    return f"{x/1e6:.0f}"


def save(fig, path, tight=True):
    os.makedirs(os.path.dirname(path), exist_ok=True)
    if tight:
        fig.tight_layout()
    fig.savefig(path, bbox_inches="tight")
    plt.close(fig)
    print(f"  -> {path}")


# ---------------------------------------------------------------------------

def read_metrics_csv(path):
    train, val = [], []
    with open(path) as f:
        for row in csv.DictReader(f):
            tok = int(row["tokens_vus"])
            if row["loss_val"]:
                val.append((tok, float(row["loss_val"])))
            elif row["loss_train"]:
                train.append((tok, float(row["loss_train"])))
    return train, val


def plot_loss(csv_path, out):
    if not os.path.exists(csv_path):
        print(f"  (loss ignorée : {csv_path} absent)")
        return
    train, val = read_metrics_csv(csv_path)
    if not train:
        print("  (loss ignorée : CSV vide)")
        return

    fig, ax = plt.subplots(figsize=(8, 5))
    xs = [t for t, _ in train]
    ys = [l for _, l in train]
    ax.plot(xs, ys, color=SERIES[0], alpha=0.85, label="entraînement")
    if val:
        vx = [t for t, _ in val]
        vy = [l for _, l in val]
        ax.plot(vx, vy, color=SERIES[1], marker="o", markersize=5,
                label="validation")

    # On réserve une marge à droite pour poser les étiquettes hors de la zone
    # tracée, plutôt que de les superposer aux courbes et à l'axe.
    xmax = max(xs)
    ax.set_xlim(min(xs) - 0.01 * xmax, xmax * 1.26)

    if val:
        # Les deux courbes finissent souvent très proches. On sépare donc les
        # étiquettes verticalement d'une fraction fixe de la hauteur du
        # graphique au lieu de les poser sur leurs valeurs respectives, qui se
        # chevaucheraient.
        span = max(max(ys), max(vy))
        ax.annotate(f"entraînement  {ys[-1]:.3f}",
                    xy=(xs[-1], ys[-1]), xytext=(10, -4),
                    textcoords="offset points", color=INK_SECONDARY,
                    fontsize=12, ha="left", va="center")
        ax.annotate(f"validation  {vy[-1]:.3f}",
                    xy=(vx[-1], vy[-1] + span * 0.06), xytext=(10, 4),
                    textcoords="offset points", color=INK_SECONDARY,
                    fontsize=12, ha="left", va="center")

    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    style_axes(ax, "La loss pendant l'entraînement",
               "Tokens vus (millions)", "Loss (entropie croisée)",
               subtitle="Plus bas = le modèle prédit mieux le coup suivant")
    if val:
        ax.legend(frameon=False, loc="upper right")
    save(fig, out)


def plot_legality(eval_glob, out, final_json=None):
    """La courbe de légalité, plus le point de mesure finale.

    Les points de la courbe viennent des évaluations intermédiaires, sur 2000
    positions chacune. La mesure publiée, elle, porte sur 20 000 positions —
    dix fois plus, donc un intervalle de confiance trois fois plus serré. On
    l'ajoute distinctement plutôt que de laisser le graphique se terminer sur
    un chiffre moins précis que celui du texte : un lecteur qui compare les
    deux ne doit pas trouver d'écart inexpliqué.
    """
    files = sorted(glob.glob(eval_glob))
    points = []
    for path in files:
        with open(path) as f:
            d = json.load(f)
        leg = d.get("legalite_generation_libre")
        if isinstance(leg, dict):
            points.append((d["tokens_vus"], leg["taux"],
                           leg["ic95_bas"], leg["ic95_haut"]))
    if not points:
        print(f"  (légalité ignorée : aucun rapport d'évaluation dans {eval_glob})")
        return
    points.sort()

    fig, ax = plt.subplots(figsize=(8, 5))
    xs = [p[0] for p in points]
    ys = [p[1] * 100 for p in points]
    lo = [p[2] * 100 for p in points]
    hi = [p[3] * 100 for p in points]

    # La bande d'incertitude est dessinée sous la courbe, en transparence :
    # elle rappelle qu'un taux mesuré sur un échantillon fini est une
    # estimation et non une vérité.
    ax.fill_between(xs, lo, hi, color=SERIES[0], alpha=0.18, linewidth=0)
    ax.plot(xs, ys, color=SERIES[0], marker="o", markersize=6)

    # L'axe vertical ne part pas de zéro, et c'est justifié : sur une courbe,
    # c'est la *position* du point qui encode la valeur, pas la longueur d'une
    # forme. Tronquer l'axe d'un diagramme en barres serait malhonnête — la
    # barre mentirait sur son rapport à ses voisines — mais ici, partir de zéro
    # écraserait toute la progression de 85 à 97 % dans le dernier dixième du
    # graphique. On borne donc juste sous le minimum observé, et le premier
    # point est étiqueté pour que le niveau de départ reste explicite.
    ax.set_ylim(min(lo) - 4, 100.8)
    ax.set_xlim(min(xs) - 0.03 * max(xs), max(xs) * 1.10)

    ax.annotate(f"{ys[0]:.1f} %", xy=(xs[0], ys[0]),
                xytext=(4, -20), textcoords="offset points",
                color=INK_SECONDARY, fontsize=13, ha="left")

    # Le point de mesure finale, sur dix fois plus de positions.
    label_final = f"{ys[-1]:.1f} %"
    if final_json and os.path.exists(final_json):
        with open(final_json) as f:
            fd = json.load(f)
        fl = fd.get("legalite_generation_libre")
        if isinstance(fl, dict):
            fx, fy = fd["tokens_vus"], fl["taux"] * 100
            ax.plot([fx], [fy], marker="*", markersize=20, color=SERIES[1],
                    linestyle="none", zorder=5)
            ax.annotate(f"{fy:.2f} %\nmesure finale\nsur {fl['positions_testees']:,}"
                        .replace(",", " ") + " positions",
                        xy=(fx, fy), xytext=(-10, -46),
                        textcoords="offset points", color=INK_PRIMARY,
                        fontsize=13, fontweight="bold", ha="right")
            label_final = None
    if label_final:
        ax.annotate(label_final, xy=(xs[-1], ys[-1]),
                    xytext=(-4, -22), textcoords="offset points",
                    color=INK_PRIMARY, fontsize=16, fontweight="bold",
                    ha="right")

    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    style_axes(ax, "Le modèle apprend les règles sans qu'on les lui donne",
               "Tokens vus (millions)", "Coups légaux (%)",
               subtitle="Génération libre, sans aucune contrainte · bande = IC 95 %")
    save(fig, out)


RULE_LABELS = {
    "roque": "Roque",
    "en_passant": "Prise en passant",
    "promotion": "Promotion",
    "echec": "Sortie d'échec",
}


def plot_rules(eval_glob, out):
    """Quatre règles, quatre panneaux.

    On pourrait superposer les quatre courbes sur un seul graphique, mais il
    faudrait alors quatre couleurs — et au-delà de deux séries, la palette
    validée ne garantit plus un écart perceptuel suffisant en vision des
    couleurs déficiente. La règle est alors de facetter plutôt que d'inventer
    des teintes : chaque panneau porte une seule courbe, dans la même couleur,
    et c'est le titre du panneau qui porte l'identité. Les quatre partagent le
    même axe vertical, ce qui est justement ce qui rend la comparaison lisible.
    """
    files = sorted(glob.glob(eval_glob))
    series = {r: [] for r in RULE_LABELS}
    for path in files:
        with open(path) as f:
            d = json.load(f)
        tests = d.get("tests_par_regle", {})
        for rule in RULE_LABELS:
            e = tests.get(rule)
            if isinstance(e, dict) and e.get("positions"):
                series[rule].append((d["tokens_vus"],
                                     e["top1_legal"]["taux"] * 100))
    if not any(series.values()):
        print(f"  (règles ignorées : aucun rapport dans {eval_glob})")
        return

    fig, axes = plt.subplots(2, 2, figsize=(9, 7), sharex=True, sharey=True)
    for ax, (rule, label) in zip(axes.flat, RULE_LABELS.items()):
        pts = sorted(series[rule])
        if not pts:
            ax.set_visible(False)
            continue
        xs = [p[0] for p in pts]
        ys = [p[1] for p in pts]
        ax.plot(xs, ys, color=SERIES[0], marker="o", markersize=4)
        ax.text(0.0, 1.04, label, transform=ax.transAxes, color=INK_PRIMARY,
                fontsize=14, fontweight="bold", va="bottom")
        ax.annotate(f"{ys[-1]:.0f} %", xy=(xs[-1], ys[-1]),
                    xytext=(-4, 8), textcoords="offset points",
                    color=INK_SECONDARY, fontsize=12, ha="right")
        # Même échelle pour les quatre panneaux, bornée juste sous la plus
        # basse valeur observée. C'est tout l'intérêt du facettage : partir de
        # zéro tasserait les quatre courbes contre le haut du cadre et rendrait
        # invisible l'écart de vingt points qui est justement le sujet.
        ax.set_ylim(70, 102)
        ax.grid(True, axis="y", linewidth=1.0, alpha=0.9)
        ax.set_axisbelow(True)
        for side in ("top", "right"):
            ax.spines[side].set_visible(False)
        for side in ("left", "bottom"):
            ax.spines[side].set_color(GRID)
        ax.xaxis.set_major_formatter(FuncFormatter(millions))

    fig.supxlabel("Tokens vus (millions)", color=INK_SECONDARY, fontsize=14)
    fig.supylabel("Coup le plus probable légal (%)", color=INK_SECONDARY,
                  fontsize=14)
    # Le titre est posé APRÈS tight_layout, sinon celui-ci recalcule les marges
    # et fait remonter le sous-titre dans les jambages du titre.
    fig.tight_layout()
    fig.suptitle("Le modèle apprend les exceptions avant la règle de base",
                 color=INK_PRIMARY, fontsize=17, fontweight="bold",
                 x=0.01, ha="left", y=1.07)
    fig.text(0.01, 1.015,
             "Part des positions où le coup le plus probable est légal",
             color=INK_SECONDARY, fontsize=12, ha="left")
    save(fig, out, tight=False)


def plot_elo(elo_json, out):
    # La courbe peut venir du rapport de l'échelle (instantanés les uns contre
    # les autres) ou du rapport général. On prend le premier qui la contient.
    points = []
    for path in (elo_json, os.path.join(os.path.dirname(elo_json),
                                        "elo_ladder.json")):
        if not os.path.exists(path):
            continue
        with open(path) as f:
            d = json.load(f)
        ech = d.get("echelle_instantanes", {})
        points = [p for p in ech.get("points", [])
                  if p.get("elo_diff") is not None]
        if points:
            break
    if not points:
        print(f"  (Elo ignoré : aucune courbe trouvée)")
        return

    fig, ax = plt.subplots(figsize=(8, 5))
    xs = [p["tokens_vus"] for p in points]
    ys = [p["elo_diff"] for p in points]

    # Barres d'erreur asymétriques. Un instantané qui ne gagne aucune partie a
    # une borne basse mathématiquement infinie : la formule d'Elo diverge quand
    # le score approche zéro. On ne peut donc pas lui donner une barre
    # symétrique sans inventer une précision qui n'existe pas. Ces points sont
    # tracés avec une flèche vers le bas, qui se lit « au moins aussi mauvais
    # que ça ».
    lo_arm, hi_arm, unbounded = [], [], []
    for p, y in zip(points, ys):
        b, h = p.get("elo_diff_ic95_bas"), p.get("elo_diff_ic95_haut")
        hi_arm.append((h - y) if h is not None else 0.0)
        if b is None:
            lo_arm.append(0.0)
            unbounded.append((p["tokens_vus"], y))
        else:
            lo_arm.append(y - b)
    err = [lo_arm, hi_arm]
    # La ligne du zéro est la référence : c'est le niveau du modèle final.
    # L'échelle est donc relative, et le dire explicitement évite qu'on lise
    # ces valeurs comme un Elo Lichess.
    ax.axhline(0, color=INK_MUTED, linewidth=1.2, linestyle="--")
    ax.errorbar(xs, ys, yerr=err, color=SERIES[0], marker="o", markersize=6,
                capsize=4, elinewidth=1.5)
    # Les points dont la borne basse est indéterminée reçoivent une flèche.
    for x, y in unbounded:
        ax.annotate("", xy=(x, y - 90), xytext=(x, y),
                    arrowprops=dict(arrowstyle="-|>", color=SERIES[0],
                                    linewidth=1.5, shrinkA=4, shrinkB=0))

    # Le libellé de la ligne de référence va à gauche, là où la courbe est
    # loin d'elle : posé à droite il chevaucherait l'étiquette du dernier
    # point, qui est justement collé au zéro.
    ax.annotate("niveau du modèle final", xy=(xs[0], 0), xytext=(6, 8),
                textcoords="offset points", color=INK_SECONDARY, fontsize=12,
                ha="left")
    ax.annotate(f"{ys[0]:+.0f}", xy=(xs[0], ys[0]), xytext=(8, 2),
                textcoords="offset points", color=INK_SECONDARY, fontsize=13)
    ax.annotate(f"{ys[-1]:+.0f}", xy=(xs[-1], ys[-1]), xytext=(-8, 6),
                textcoords="offset points", color=INK_PRIMARY, fontsize=14,
                fontweight="bold", ha="right")
    ax.set_xlim(min(xs) - 0.05 * max(xs), max(xs) * 1.08)
    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    span = ys[-1] - ys[0]
    style_axes(ax, f"{span:.0f} points d'Elo gagnés en deux heures",
               "Tokens vus (millions)", "Écart d'Elo au modèle final",
               subtitle="Chaque instantané joue 60 parties contre le modèle final · "
                        "barres = IC 95 %, flèche = borne basse indéterminée")
    save(fig, out)


def plot_comparaison_loss(csv_a, csv_b, label_a, label_b, out):
    """Les deux entraînements sur le même axe.

    C'est le graphique qui raconte le mieux l'effet de l'échelle : même tâche,
    même code, seuls la taille du modèle et le volume de données changent. On
    trace en fonction des tokens vus et non des heures écoulées, parce que le
    temps mélangerait deux effets — le modèle plus gros voit moins de tokens
    par seconde, ce qui n'a rien à voir avec ce qu'il apprend.
    """
    series = []
    for path, label in ((csv_a, label_a), (csv_b, label_b)):
        if not os.path.exists(path):
            print(f"  (comparaison : {path} absent)")
            return
        _, val = read_metrics_csv(path)
        if not val:
            print(f"  (comparaison : pas de loss de validation dans {path})")
            return
        series.append((label, val))

    fig, ax = plt.subplots(figsize=(8, 5))
    for i, (label, val) in enumerate(series):
        xs = [t for t, _ in val]
        ys = [l for _, l in val]
        ax.plot(xs, ys, color=SERIES[i], label=label, linewidth=2)
        ax.annotate(f"{label}  {ys[-1]:.3f}", xy=(xs[-1], ys[-1]),
                    xytext=(8, 0), textcoords="offset points",
                    color=INK_SECONDARY, fontsize=12, va="center")

    xmax = max(max(t for t, _ in v) for _, v in series)
    ax.set_xlim(0, xmax * 1.32)
    lo = min(min(l for _, l in v) for _, v in series)
    ax.set_ylim(lo * 0.93, 3.2)
    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    style_axes(ax, "Ce que change un modèle trois fois plus gros",
               "Tokens vus (millions)", "Loss de validation",
               subtitle="Même code, même tâche, seules la taille et la "
                        "quantité de données changent")
    ax.legend(frameon=False, loc="upper right")
    save(fig, out)


def plot_comparaison_legalite(glob_a, glob_b, label_a, label_b, out):
    """Le taux de coups légaux des deux runs, superposé."""
    series = []
    for pattern, label in ((glob_a, label_a), (glob_b, label_b)):
        pts = []
        for path in sorted(glob.glob(pattern)):
            with open(path) as f:
                d = json.load(f)
            leg = d.get("legalite_generation_libre")
            if isinstance(leg, dict):
                pts.append((d["tokens_vus"], leg["taux"] * 100))
        if pts:
            series.append((label, sorted(pts)))
    if len(series) < 2:
        print("  (comparaison légalité : il manque un des deux runs)")
        return

    fig, ax = plt.subplots(figsize=(8, 5))
    for i, (label, pts) in enumerate(series):
        xs = [p[0] for p in pts]
        ys = [p[1] for p in pts]
        ax.plot(xs, ys, color=SERIES[i], marker="o", markersize=5, label=label)
        ax.annotate(f"{ys[-1]:.1f} %", xy=(xs[-1], ys[-1]),
                    xytext=(8, 0), textcoords="offset points",
                    color=INK_SECONDARY, fontsize=12, va="center")

    xmax = max(max(p[0] for p in pts) for _, pts in series)
    ax.set_xlim(0, xmax * 1.18)
    ymin = min(min(p[1] for p in pts) for _, pts in series)
    ax.set_ylim(ymin - 3, 100.8)
    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    style_axes(ax, "Le taux de coups légaux, à deux échelles",
               "Tokens vus (millions)", "Coups légaux (%)",
               subtitle="Génération libre, sans aucune contrainte")
    ax.legend(frameon=False, loc="lower right")
    save(fig, out)


def plot_comparaison_accord(glob_a, glob_b, label_a, label_b, out):
    """L'accord avec le coup humain, pour les deux runs.

    C'est le pendant du graphique de légalité, et il raconte l'inverse : là où
    les deux courbes de légalité se superposent, celles-ci se séparent
    nettement. Les deux graphiques n'ont de sens que côte à côte — c'est leur
    contraste qui dit ce que la taille du modèle achète réellement.
    """
    series = []
    for pattern, label in ((glob_a, label_a), (glob_b, label_b)):
        pts = []
        for path in sorted(glob.glob(pattern)):
            with open(path) as f:
                d = json.load(f)
            a = d.get("accord_humain")
            if isinstance(a, dict):
                pts.append((d["tokens_vus"], a["top1"]["taux"] * 100))
        if pts:
            series.append((label, sorted(pts)))
    if len(series) < 2:
        print("  (comparaison accord : il manque un des deux runs)")
        return

    fig, ax = plt.subplots(figsize=(8, 5))
    for i, (label, pts) in enumerate(series):
        xs = [p[0] for p in pts]
        ys = [p[1] for p in pts]
        ax.plot(xs, ys, color=SERIES[i], marker="o", markersize=5, label=label)
        ax.annotate(f"{ys[-1]:.1f} %", xy=(xs[-1], ys[-1]),
                    xytext=(8, 0), textcoords="offset points",
                    color=INK_SECONDARY, fontsize=12, va="center")

    xmax = max(max(p[0] for p in pts) for _, pts in series)
    ax.set_xlim(0, xmax * 1.18)
    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    style_axes(ax, "Là où la taille du modèle change quelque chose",
               "Tokens vus (millions)", "Coup humain prédit (%)",
               subtitle="Part des positions où le coup le plus probable est "
                        "celui joué par l'humain")
    ax.legend(frameon=False, loc="lower right")
    save(fig, out)


def plot_lengths(stats_json, out):
    if not os.path.exists(stats_json):
        print(f"  (longueurs ignorées : {stats_json} absent)")
        return
    with open(stats_json) as f:
        d = json.load(f)
    hist = d.get("histogramme_longueurs")
    if not hist:
        print("  (longueurs ignorées : pas d'histogramme)")
        return

    lengths = sorted((int(k), v) for k, v in hist.items())
    xs = [k for k, _ in lengths]
    ys = [v for _, v in lengths]

    fig, ax = plt.subplots(figsize=(8, 5))
    ax.bar(xs, ys, width=1.0, color=SERIES[0], linewidth=0)
    total = sum(ys)
    mean = sum(k * v for k, v in lengths) / total
    ax.axvline(mean, color=SERIES[1], linewidth=2, linestyle="--")
    ax.annotate(f"moyenne\n{mean:.0f} demi-coups", xy=(mean, max(ys) * 0.82),
                xytext=(10, 0), textcoords="offset points",
                color=INK_SECONDARY, fontsize=12)
    style_axes(ax, "Longueur des parties retenues",
               "Demi-coups", "Nombre de parties",
               subtitle=f"{total:,} parties après filtrage".replace(",", " "))
    save(fig, out)


def plot_throughput(csv_path, out):
    if not os.path.exists(csv_path):
        return
    xs, ys = [], []
    with open(csv_path) as f:
        for row in csv.DictReader(f):
            if row["mfu"]:
                xs.append(int(row["tokens_vus"]))
                ys.append(float(row["mfu"]) * 100)
    if not xs:
        return
    fig, ax = plt.subplots(figsize=(8, 5))
    ax.plot(xs, ys, color=SERIES[0], alpha=0.9)
    median = sorted(ys)[len(ys) // 2]
    ax.annotate(f"médiane {median:.0f} %", xy=(xs[len(xs) // 2], median),
                xytext=(0, 14), textcoords="offset points",
                color=INK_PRIMARY, fontsize=14, fontweight="bold", ha="center")
    ax.xaxis.set_major_formatter(FuncFormatter(millions))
    ax.set_ylim(0, 100)
    style_axes(ax, "Utilisation réelle du GPU",
               "Tokens vus (millions)", "MFU (%)",
               subtitle="Part de la puissance mesurée de la carte réellement exploitée")
    save(fig, out)


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--run-name", default="run1")
    p.add_argument("--log-dir", default="logs")
    p.add_argument("--out-dir", default="figures")
    p.add_argument("--suffixe", default="",
                   help="suffixe des noms de fichiers, ex. _run2")
    p.add_argument("--eval-glob", default="eval_step*.json")
    p.add_argument("--elo-json", default="elo_report.json")
    p.add_argument("--comparer", nargs=2, metavar=("RUN_A", "RUN_B"),
                   default=None,
                   help="produit aussi les graphiques comparant deux runs")
    args = p.parse_args()
    sfx = args.suffixe

    print("Génération des graphiques...")
    plot_loss(os.path.join(args.log_dir, f"{args.run_name}_metrics.csv"),
              os.path.join(args.out_dir, f"loss{sfx}.png"))
    plot_legality(os.path.join(args.log_dir, args.eval_glob),
                  os.path.join(args.out_dir, f"legalite{sfx}.png"),
                  os.path.join(args.log_dir,
                               f"eval_final{sfx}.json" if sfx
                               else "eval_final.json"))
    plot_rules(os.path.join(args.log_dir, args.eval_glob),
               os.path.join(args.out_dir, f"regles{sfx}.png"))
    plot_elo(os.path.join(args.log_dir, args.elo_json),
             os.path.join(args.out_dir, f"elo{sfx}.png"))
    plot_lengths(os.path.join(args.log_dir, "phase1_parse_stats.json"),
                 os.path.join(args.out_dir, "longueurs_parties.png"))
    plot_throughput(os.path.join(args.log_dir, f"{args.run_name}_metrics.csv"),
                    os.path.join(args.out_dir, f"mfu{sfx}.png"))

    if args.comparer:
        a, b = args.comparer
        la = "51 M param., 1 mois" if a == "run1" else a
        lb = "142 M param., 4 mois" if b == "run2" else b
        plot_comparaison_loss(
            os.path.join(args.log_dir, f"{a}_metrics.csv"),
            os.path.join(args.log_dir, f"{b}_metrics.csv"),
            la, lb, os.path.join(args.out_dir, "comparaison_loss.png"))
        plot_comparaison_legalite(
            os.path.join(args.log_dir, "eval_step*.json"),
            os.path.join(args.log_dir, "eval_run2_step*.json"),
            la, lb, os.path.join(args.out_dir, "comparaison_legalite.png"))
        plot_comparaison_accord(
            os.path.join(args.log_dir, "eval_step*.json"),
            os.path.join(args.log_dir, "eval_run2_step*.json"),
            la, lb, os.path.join(args.out_dir, "comparaison_accord.png"))


if __name__ == "__main__":
    main()
plots_article.py314 lines

The one-off visuals: Sicilian, temperature, filtering, error patterns.

"""Graphiques supplémentaires, pensés pour un article grand public.

Ceux de `plots.py` suivent la progression de l'entraînement. Ceux-ci illustrent
des constats ponctuels — ils répondent chacun à une question qu'un lecteur se
pose, et fonctionnent isolément.

Règle de rendu appliquée partout : **les diagrammes en barres partent de zéro**,
contrairement aux courbes. La différence n'est pas cosmétique. Sur une barre,
c'est la *longueur* qui encode la valeur ; tronquer l'axe fait mentir le rapport
visuel entre deux barres. Sur une courbe, c'est la *position* du point, et un
axe tronqué reste honnête tant qu'il est lisiblement gradué.

Usage :
    python plots_article.py
"""

import json
import os

import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

from plots import (SURFACE, INK_PRIMARY, INK_SECONDARY, INK_MUTED, GRID,
                   SERIES, save, style_axes)

OUT = "figures"


def barres_horizontales(ax, labels, valeurs, couleurs=None, suffixe=" %",
                        fmt="{:.1f}"):
    """Barres horizontales, valeur écrite au bout de chaque barre.

    L'horizontale est préférée quand les étiquettes sont des mots : elles se
    lisent normalement, sans rotation à 45 degrés qui oblige le lecteur à
    pencher la tête — un détail qui compte beaucoup sur mobile.
    """
    y = range(len(labels))
    couleurs = couleurs or [SERIES[0]] * len(labels)
    ax.barh(list(y), valeurs, color=couleurs, height=0.62)
    ax.set_yticks(list(y))
    ax.set_yticklabels(labels, fontsize=13)
    ax.invert_yaxis()
    vmax = max(valeurs)
    for i, v in enumerate(valeurs):
        ax.text(v + vmax * 0.015, i, fmt.format(v) + suffixe,
                va="center", ha="left", fontsize=13, color=INK_PRIMARY,
                fontweight="bold")
    ax.set_xlim(0, vmax * 1.22)
    ax.grid(True, axis="x", linewidth=1.0, alpha=0.9)
    ax.set_axisbelow(True)
    for side in ("top", "right", "left"):
        ax.spines[side].set_visible(False)
    ax.spines["bottom"].set_color(GRID)
    ax.tick_params(axis="y", length=0)


# ---------------------------------------------------------------------------

def g_sicilienne():
    """Ce que le modèle propose après 1.e4 c5.

    Mesuré avec engine.py sur le checkpoint run1_best, probabilités
    renormalisées sur les seuls coups légaux.
    """
    coups = ["Cf3 : Sicilienne ouverte", "Cc3 : variante fermée",
             "c3 : Alapine", "d4 : poussée centrale",
             "Fc4 : fou italien", "f4 : attaque Grand Prix"]
    p = [54.7, 10.8, 8.9, 7.4, 5.8, 4.2]
    fig, ax = plt.subplots(figsize=(8.5, 4.6))
    barres_horizontales(ax, coups, p)
    ax.set_xlabel("Probabilité attribuée par le modèle (%)")
    # On abaisse le haut de l'axe pour dégager la place du titre et du
    # sous-titre, sinon ce dernier chevauche la première barre.
    fig.subplots_adjust(top=0.76, bottom=0.14, right=0.97)

    # Titre et sous-titre centrés sur le CONTENU (libellés y + barres), pas sur
    # l'axe. Sinon, comme les libellés débordent loin à gauche et que save()
    # recadre en bbox tight, le titre paraît décalé à droite. On calcule le
    # milieu réel au rendu.
    fig.canvas.draw()
    rend = fig.canvas.get_renderer()
    gauche = min(l.get_window_extent(rend).x0 for l in ax.get_yticklabels())
    droite = ax.get_window_extent(rend).x1
    milieu = fig.transFigure.inverted().transform(((gauche + droite) / 2, 0))[0]
    fig.text(milieu, 0.95, "Il connaît la théorie des ouvertures",
             ha="center", color=INK_PRIMARY, fontsize=17, fontweight="bold")
    fig.text(milieu, 0.865,
             "Après 1.e4 c5 : les six grandes réponses à la Sicilienne, "
             "dans leur ordre de popularité réel",
             ha="center", color=INK_SECONDARY, fontsize=12)
    # tight=False : sinon fig.tight_layout() annule le subplots_adjust ci-dessus
    # et fait remonter l'axe sous le sous-titre.
    save(fig, os.path.join(OUT, "sicilienne.png"), tight=False)


def g_temperature():
    """L'effet du réglage d'échantillonnage sur le comportement observé."""
    temps = [1.0, 0.8, 0.6, 0.4, 0.2]
    sans_faute = [52.5, 68.0, 80.5, 82.5, 87.0]
    longueur = [48.2, 54.5, 64.1, 63.6, 60.5]

    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4.4))
    for ax, vals, titre, ylab, ref in (
            (ax1, sans_faute, "Parties sans coup illégal", "%", None),
            (ax2, longueur, "Longueur des parties", "demi-coups", 71.6)):
        ax.plot(temps, vals, color=SERIES[0], marker="o", markersize=7)
        ax.invert_xaxis()
        if ref:
            ax.axhline(ref, color=SERIES[1], linewidth=2, linestyle="--")
            # Sous la ligne et non au-dessus : au-dessus elle percutait le
            # titre du panneau, la référence étant proche du haut du cadre.
            ax.text(0.98, ref - 2.2, f"vraies parties : {ref}",
                    color=INK_SECONDARY, fontsize=11, ha="left")
        ax.text(0.0, 1.05, titre, transform=ax.transAxes, color=INK_PRIMARY,
                fontsize=14, fontweight="bold", va="bottom")
        ax.set_xlabel("Température d'échantillonnage")
        ax.set_ylabel(ylab)
        ax.grid(True, axis="y", linewidth=1.0, alpha=0.9)
        ax.set_axisbelow(True)
        for s in ("top", "right"):
            ax.spines[s].set_visible(False)
        for s in ("left", "bottom"):
            ax.spines[s].set_color(GRID)

    fig.tight_layout()
    fig.suptitle("Le même modèle, mesuré à cinq réglages",
                 color=INK_PRIMARY, fontsize=17, fontweight="bold",
                 x=0.01, ha="left", y=1.16)
    fig.text(0.01, 1.045,
             "La température n'est pas un paramètre du modèle mais de la "
             "mesure. Changer d'instrument change le résultat",
             color=INK_SECONDARY, fontsize=12, ha="left")
    save(fig, os.path.join(OUT, "temperature.png"), tight=False)


def g_motifs_echec():
    """De quoi sont faites les 2 % d'erreurs."""
    try:
        d = json.load(open("logs/eval_final.json"))
        m = d["legalite_generation_libre"]["motifs_echec"]
    except Exception:
        print("  (motifs : eval_final.json absent)")
        return
    noms = {"deplacement_illegal": "Coup impossible\ndans cette position",
            "token_special": "Symbole de service\nproposé comme un coup",
            "case_depart_vide": "Déplace une pièce\nqui n'existe pas",
            "piece_adverse": "Déplace une pièce\nde l'adversaire"}
    items = sorted(m.items(), key=lambda kv: -kv[1])
    labels = [noms.get(k, k) for k, _ in items]
    vals = [v for _, v in items]

    fig, ax = plt.subplots(figsize=(8.5, 4.2))
    barres_horizontales(ax, labels, vals, suffixe="", fmt="{:.0f}")
    ax.set_xlabel("Occurrences sur 20 000 coups générés")
    ax.text(0.0, 1.16, "Quand il se trompe, il se trompe finement",
            transform=ax.transAxes, color=INK_PRIMARY, fontsize=17,
            fontweight="bold", va="bottom")
    ax.text(0.0, 1.04,
            "Dans la moitié des cas, le coup proposé est géométriquement "
            "correct, il est juste impossible ici",
            transform=ax.transAxes, color=INK_SECONDARY, fontsize=12,
            va="bottom")
    save(fig, os.path.join(OUT, "motifs_echec.png"))


def g_filtrage():
    """Ce que le filtrage jette, à l'échelle d'un mois de Lichess."""
    try:
        d = json.load(open("logs/phase1_parse_stats.json"))
        r = d["rejets"]
        gardees = d["parties_conservees"]
    except Exception:
        print("  (filtrage : phase1_parse_stats.json absent)")
        return
    noms = {"elo_hors_bornes": "Elo hors de 1800-2600",
            "cadence_exclue": "Bullet et ultrabullet",
            "terminaison": "Fin anormale (déconnexion…)",
            "trop_courte": "Moins de 20 demi-coups",
            "trop_longue": "Plus de 300 demi-coups"}
    items = sorted(r.items(), key=lambda kv: -kv[1])
    labels = ["CONSERVÉES"] + [noms.get(k, k) for k, _ in items]
    vals = [gardees / 1e6] + [v / 1e6 for _, v in items]
    couleurs = [SERIES[1]] + [SERIES[0]] * len(items)

    fig, ax = plt.subplots(figsize=(8.5, 4.6))
    barres_horizontales(ax, labels, vals, couleurs, suffixe=" M", fmt="{:.1f}")
    ax.set_xlabel("Millions de parties, sur un mois de Lichess")
    ax.text(0.0, 1.16, "Sur 89 millions de parties, on en garde 11",
            transform=ax.transAxes, color=INK_PRIMARY, fontsize=17,
            fontweight="bold", va="bottom")
    ax.text(0.0, 1.04,
            "Plus d'un tiers des parties jouées sur Lichess sont du bullet",
            transform=ax.transAxes, color=INK_SECONDARY, fontsize=12,
            va="bottom")
    save(fig, os.path.join(OUT, "filtrage.png"))


def g_bridage_gpu():
    """Le GPU bridé : 110 watts qui valent un facteur 2.5."""
    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.5, 4.2))
    for ax, vals, titre, ylab in (
            (ax1, [24.98, 62.86], "Calcul délivré", "TFLOPS bf16"),
            (ax2, [596, 1506], "Fréquence du GPU", "MHz")):
        b = ax.bar(["bridé\n170 W", "libéré\n280 W"], vals,
                   color=[INK_MUTED, SERIES[0]], width=0.55)
        for rect, v in zip(b, vals):
            ax.text(rect.get_x() + rect.get_width() / 2, v * 1.03,
                    f"{v:,.0f}".replace(",", " ") if v > 100 else f"{v:.2f}",
                    ha="center", fontsize=14, fontweight="bold",
                    color=INK_PRIMARY)
        ax.text(0.0, 1.05, titre, transform=ax.transAxes, color=INK_PRIMARY,
                fontsize=14, fontweight="bold", va="bottom")
        ax.set_ylabel(ylab)
        ax.set_ylim(0, max(vals) * 1.2)
        ax.grid(True, axis="y", linewidth=1.0, alpha=0.9)
        ax.set_axisbelow(True)
        for s in ("top", "right"):
            ax.spines[s].set_visible(False)
        for s in ("left", "bottom"):
            ax.spines[s].set_color(GRID)

    fig.tight_layout()
    fig.suptitle("110 watts qui valent un facteur 2.5",
                 color=INK_PRIMARY, fontsize=17, fontweight="bold",
                 x=0.01, ha="left", y=1.17)
    fig.text(0.01, 1.05,
             "La carte était bridée à 40 % de sa puissance, et produisait "
             "moins qu'un GPU trois fois plus petit",
             color=INK_SECONDARY, fontsize=12, ha="left")
    save(fig, os.path.join(OUT, "bridage_gpu.png"), tight=False)


def g_regles_final():
    """Les quatre règles, en barres, sur le modèle final."""
    try:
        d = json.load(open("logs/eval_final.json"))["tests_par_regle"]
    except Exception:
        print("  (règles finales : eval_final.json absent)")
        return
    noms = [("roque", "Roque"), ("en_passant", "Prise en passant"),
            ("promotion", "Promotion"), ("echec", "Sortie d'échec")]
    labels = [n for _, n in noms]
    vals = [d[k]["top1_legal"]["taux"] * 100 for k, _ in noms]
    couleurs = [SERIES[0]] * 3 + [SERIES[1]]

    fig, ax = plt.subplots(figsize=(8.5, 3.8))
    barres_horizontales(ax, labels, vals, couleurs)
    ax.set_xlim(0, 112)
    ax.set_xlabel("Coup le plus probable légal (%), sur 500 positions")
    ax.text(0.0, 1.18, "Il apprend les exceptions avant la règle de base",
            transform=ax.transAxes, color=INK_PRIMARY, fontsize=17,
            fontweight="bold", va="bottom")
    ax.text(0.0, 1.05,
            "Sortir d'échec est la contrainte la plus fondamentale du jeu, "
            "c'est celle qu'il rate le plus",
            transform=ax.transAxes, color=INK_SECONDARY, fontsize=12,
            va="bottom")
    save(fig, os.path.join(OUT, "regles_final.png"))


def g_qwen():
    """Le chiffre central comparé, et l'effet du protocole sur la mesure.

    Quatre barres, une seule mesure : une seule couleur suffit pour nos
    modèles, une seconde distingue l'adversaire. L'axe part de zéro, comme tout
    diagramme en barres.

    Les deux barres de l'adversaire ne sont pas redondantes, elles sont le
    sujet : mesuré sur les seules positions d'ouverture il obtient 76 %, mesuré
    sur des parties menées jusqu'au mat il tombe à 36 %. Le protocole change ce
    qu'on mesure, et le montrer vaut mieux que de choisir le chiffre qui
    arrange.
    """
    import json as _j
    try:
        q_long = _j.load(open("logs/duel_qwen.json"))["legalite_premiere_tentative"]["taux"] * 100
        q_court = _j.load(open("logs/duel_qwen_run2_strict.json"))["legalite_premiere_tentative"]["taux"] * 100
        r1 = _j.load(open("logs/eval_final.json"))["legalite_generation_libre"]["taux"] * 100
        r2 = _j.load(open("logs/eval_final_run2_t10.json"))["legalite_generation_libre"]["taux"] * 100
    except Exception as e:
        print(f"  (qwen : donnees manquantes, {e})"); return

    labels = ["Mon modèle, 142 M",
              "Mon modèle, 51 M",
              "Qwen 35 Md\nsur les ouvertures",
              "Qwen 35 Md\nparties completes"]
    vals = [r2, r1, q_court, q_long]
    couleurs = [SERIES[0], SERIES[0], SERIES[1], SERIES[1]]

    fig, ax = plt.subplots(figsize=(8.5, 4.6))
    barres_horizontales(ax, labels, vals, couleurs)
    ax.set_xlim(0, 118)
    ax.set_xlabel("Coups légaux en génération libre (%)")
    ax.text(0.0, 1.20, "250 fois plus gros, et trois fois plus d'erreurs",
            transform=ax.transAxes, color=INK_PRIMARY, fontsize=17,
            fontweight="bold", va="bottom")
    ax.text(0.0, 1.06,
            "Même question posée aux deux : un coup, sans liste des coups "
            "légaux, sans seconde chance",
            transform=ax.transAxes, color=INK_SECONDARY, fontsize=12,
            va="bottom")
    save(fig, os.path.join(OUT, "qwen_legalite.png"))


if __name__ == "__main__":
    os.makedirs(OUT, exist_ok=True)
    print("Graphiques d'article :")
    for f in (g_sicilienne, g_temperature, g_motifs_echec, g_filtrage,
              g_bridage_gpu, g_regles_final, g_qwen):
        try:
            f()
        except Exception as e:
            print(f"  ECHEC {f.__name__} : {e}")
plots_hero.py421 lines

The opening cards for social networks.

"""Images d'ouverture pour les réseaux sociaux.

Différence de nature avec `plots.py` et `plots_article.py` : celles-ci sont des
images de **première lecture**. Elles doivent être comprises en vignette, à
cent pixels de côté, par quelqu'un qui fait défiler son fil sans s'arrêter.

Un graphique ne remplit jamais ce rôle. Ses axes, ses graduations et ses
annotations disparaissent à cette taille. Le graphique reste utile en seconde
image, pour celui qui a ouvert le post.

Ce qui fonctionne en vignette : du très gros texte, un seul message, et un
contraste fort. Rien d'autre.

Usage :
    python plots_hero.py
"""

import os

import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

from plots import SURFACE, INK_PRIMARY, INK_SECONDARY, INK_MUTED, SERIES, save

OUT = "figures"
# Format carré : LinkedIn recadre les images larges en vignette, et un carré
# survit au recadrage sur toutes les tailles d'écran.
TAILLE = (7.2, 7.2)


def espace(txt):
    """Espacement des lettres, que matplotlib ne sait pas faire nativement.

    On intercale une espace fine insécable entre les caractères. C'est le seul
    moyen d'obtenir l'aération d'un intertitre en petites capitales.
    """
    return "\u2009".join(txt)


def cadre(ax):
    ax.set_xlim(0, 1); ax.set_ylim(0, 1); ax.axis("off")


def hero_legalite():
    """L'image d'ouverture du post 1.

    Trois blocs : ce que le modèle reçoit, le résultat en très gros, ce que ça
    signifie. Le chiffre occupe le tiers central pour rester lisible même
    fortement réduit.
    """
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("TOUT CE QUE LE MODÈLE REÇOIT"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    # La séquence de coups, en chasse fixe : c'est littéralement l'entrée.
    # Quatre coups suffisent, et le texte tient dans le cadre. Avec cinq, le
    # point d'interrogation débordait à droite en chasse fixe.
    ax.text(0.5, 0.845, "e2e4  e7e5  g1f3  b8c6  →  ?",
            ha="center", color=INK_PRIMARY, fontsize=20,
            family="DejaVu Sans Mono")

    ax.plot([0.12, 0.88], [0.78, 0.78], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.60, "Aucun échiquier.", ha="center",
            color=INK_SECONDARY, fontsize=22)
    ax.text(0.5, 0.535, "Aucune règle du jeu.", ha="center",
            color=INK_SECONDARY, fontsize=22)

    ax.text(0.5, 0.315, "97,86 %", ha="center", color=SERIES[0],
            fontsize=88, fontweight="bold")
    ax.text(0.5, 0.235, "des coups qu'il propose sont légaux",
            ha="center", color=INK_PRIMARY, fontsize=21)

    ax.plot([0.12, 0.88], [0.17, 0.17], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.105, "Il l'a déduit seul de 800 millions de coups",
            ha="center", color=INK_SECONDARY, fontsize=17)
    ax.text(0.5, 0.045, "51 M de paramètres  ·  2 heures  ·  1 RTX 3090",
            ha="center", color=INK_MUTED, fontsize=15)

    save(fig, os.path.join(OUT, "hero_post1.png"), tight=False)


def hero_qwen():
    """Image d'ouverture du post sur le duel contre le généraliste."""
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("MÊME QUESTION, MÊMES CONDITIONS"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.855, "« Quel est le coup suivant ? »", ha="center",
            color=INK_PRIMARY, fontsize=21)

    for y, titre, taille, valeur, coul in (
            (0.63, "Mon modèle", "142 millions de paramètres", "98,85 %", SERIES[0]),
            (0.28, "Qwen3.6", "35 milliards de paramètres", "35 %", SERIES[1])):
        ax.text(0.5, y + 0.115, titre, ha="center", color=INK_PRIMARY,
                fontsize=23, fontweight="bold")
        ax.text(0.5, y + 0.065, taille, ha="center", color=INK_SECONDARY,
                fontsize=17)
        ax.text(0.5, y - 0.035, valeur, ha="center", color=coul,
                fontsize=66, fontweight="bold")

    ax.text(0.5, 0.055, "de coups légaux, sans seconde chance",
            ha="center", color=INK_MUTED, fontsize=16)

    save(fig, os.path.join(OUT, "hero_qwen.png"), tight=False)


def hero_regles():
    """Image d'ouverture du post sur l'ordre d'apprentissage.

    Recadrée « IA » et non « échecs » : le message est que, dans un même
    réseau et à données égales, certaines compétences s'acquièrent d'emblée et
    d'autres restent à la traîne. On le montre par deux nombres comparables,
    mesurés au même instant (49 M d'exemples vus) : un motif de surface est
    déjà maîtrisé quand une compétence qui exige de reconstruire un état est
    encore vingt points derrière. Chaque nombre est centré sur son bloc, ce qui
    supprime tout risque de chevauchement libellé/valeur.
    """
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.905, espace("TOUT NE S'APPREND PAS"), ha="center",
            color=INK_MUTED, fontsize=16, fontweight="bold")
    ax.text(0.5, 0.845, espace("À LA MÊME VITESSE"), ha="center",
            color=INK_MUTED, fontsize=16, fontweight="bold")

    # Deux régimes contrastés, chacun en trois lignes centrées (titre, nature,
    # nombre). Le nombre est la vedette ; les deux sont dans la même unité et
    # au même instant, donc l'écart se lit sans légende.
    blocs = [
        (0.735, 0.545, "Un motif de surface", "fréquent, toujours identique",
         "98,7 %", SERIES[0]),
        (0.410, 0.220, "Un état à reconstruire", "à déduire de la séquence",
         "76,7 %", SERIES[1]),
    ]
    for y_titre, y_val, titre, nature, val, coul in blocs:
        ax.text(0.5, y_titre, titre, ha="center", color=INK_PRIMARY,
                fontsize=25, fontweight="bold")
        ax.text(0.5, y_titre - 0.052, nature, ha="center", color=INK_SECONDARY,
                fontsize=16)
        ax.text(0.5, y_val, val, ha="center", va="center", color=coul,
                fontsize=54, fontweight="bold")

    ax.plot([0.1, 0.9], [0.135, 0.135], color="#e6e5e1", linewidth=2)
    ax.text(0.5, 0.08, "Réponses valides après les mêmes 49 M d'exemples",
            ha="center", color=INK_SECONDARY, fontsize=16)
    ax.text(0.5, 0.033, "Combler l'écart lui coûte près de 10× plus de données",
            ha="center", color=INK_MUTED, fontsize=14.5)

    save(fig, os.path.join(OUT, "hero_regles.png"), tight=False)



def hero_sicilienne():
    """Image d'ouverture du post sur les ouvertures, cadrée « émergence ».

    Le message n'est pas « le modèle joue bien », c'est qu'une hiérarchie de
    préférences a émergé de la seule prédiction du coup suivant. Structure
    reprise de hero_legalite (révélation entrée -> chiffre -> sens), qui reste
    lisible en vignette. Le nombre est centré, donc aucun chevauchement.
    """
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.92, espace("ON LUI DONNE DEUX COUPS"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.85, "e2e4   c7c5   →   ?", ha="center",
            color=INK_PRIMARY, fontsize=21, family="DejaVu Sans Mono")

    ax.plot([0.12, 0.88], [0.785, 0.785], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.66, "Il mise", ha="center",
            color=INK_SECONDARY, fontsize=21)
    ax.text(0.5, 0.52, "54,7 %", ha="center", va="center", color=SERIES[0],
            fontsize=74, fontweight="bold")
    ax.text(0.5, 0.385, "sur le coup le plus joué au monde", ha="center",
            color=INK_PRIMARY, fontsize=21)

    ax.plot([0.12, 0.88], [0.315, 0.315], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.225, "Ses six réponses suivent la popularité réelle,",
            ha="center", color=INK_SECONDARY, fontsize=18)
    ax.text(0.5, 0.165, "que personne ne lui a jamais donnée.", ha="center",
            color=INK_SECONDARY, fontsize=18)

    save(fig, os.path.join(OUT, "hero_sicilienne.png"), tight=False)


def hero_overfit():
    """Post technique : le test de surapprentissage.

    Une courbe minuscule ne dirait rien en vignette. On montre quatre paliers
    de la descente, en très gros, comme un compte à rebours.
    """
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("AVANT DE LANCER 2 HEURES"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.865, espace("JE VÉRIFIE EN 2 MINUTES"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    ax.text(0.5, 0.755, "Le modèle sait-il apprendre", ha="center",
            color=INK_PRIMARY, fontsize=22)
    ax.text(0.5, 0.700, "100 parties par cœur ?", ha="center",
            color=INK_PRIMARY, fontsize=22)

    paliers = [("step 10", "7,13", INK_MUTED),
               ("step 50", "2,28", INK_MUTED),
               ("step 100", "0,16", SERIES[0]),
               ("step 270", "0,035", SERIES[0])]
    y = 0.545
    for etape, val, coul in paliers:
        ax.text(0.30, y, etape, ha="right", va="center",
                color=INK_SECONDARY, fontsize=20)
        ax.text(0.40, y, val, ha="left", va="center",
                color=coul, fontsize=36, fontweight="bold")
        y -= 0.105

    ax.plot([0.12, 0.88], [0.185, 0.185], color="#e6e5e1", linewidth=2)
    ax.text(0.5, 0.115, "Si la loss ne tombe pas, il y a un bug.", ha="center",
            color=INK_SECONDARY, fontsize=19)
    ax.text(0.5, 0.060, "Et le run complet le reproduira toute la nuit.",
            ha="center", color=INK_SECONDARY, fontsize=19)

    save(fig, os.path.join(OUT, "hero_overfit.png"), tight=False)


def hero_loss_piege():
    """Post technique : la loss anormalement basse."""
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("SUR UN MODÈLE"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.868, espace("QUI N'AVAIT RIEN APPRIS"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    ax.text(0.26, 0.700, "loss attendue", ha="center",
            color=INK_SECONDARY, fontsize=18)
    ax.text(0.26, 0.600, "7,586", ha="center",
            color=INK_MUTED, fontsize=44, fontweight="bold")
    ax.text(0.26, 0.535, "le pur hasard", ha="center",
            color=INK_MUTED, fontsize=15)

    ax.text(0.74, 0.700, "loss mesurée", ha="center",
            color=INK_SECONDARY, fontsize=18)
    ax.text(0.74, 0.600, "5,998", ha="center",
            color=SERIES[1], fontsize=44, fontweight="bold")
    ax.text(0.74, 0.535, "bien meilleure", ha="center",
            color=SERIES[1], fontsize=15)

    ax.plot([0.5, 0.5], [0.50, 0.76], color="#e6e5e1", linewidth=2)
    ax.plot([0.10, 0.90], [0.435, 0.435], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.335, "Une loss trop belle", ha="center",
            color=INK_PRIMARY, fontsize=24, fontweight="bold")
    ax.text(0.5, 0.255, "n'est jamais une bonne nouvelle.", ha="center",
            color=INK_PRIMARY, fontsize=24, fontweight="bold")

    ax.text(0.5, 0.145, "C'est une piste de fuite de données", ha="center",
            color=INK_SECONDARY, fontsize=19)
    ax.text(0.5, 0.090, "à remonter.", ha="center",
            color=INK_SECONDARY, fontsize=19)

    save(fig, os.path.join(OUT, "hero_loss_piege.png"), tight=False)


def hero_chinchilla():
    """Post technique : combien de données pour quelle taille."""
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("COMBIEN DE DONNÉES"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.868, espace("POUR QUELLE TAILLE ?"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    ax.text(0.5, 0.700, "20", ha="center", color=SERIES[0],
            fontsize=110, fontweight="bold")
    ax.text(0.5, 0.615, "tokens de données", ha="center",
            color=INK_PRIMARY, fontsize=24)
    ax.text(0.5, 0.560, "par paramètre", ha="center",
            color=INK_PRIMARY, fontsize=24)

    ax.plot([0.12, 0.88], [0.485, 0.485], color="#e6e5e1", linewidth=2)

    for y, nom, val in ((0.400, "Mon 1er modèle, 51 M", "15,7"),
                        (0.300, "Mon 2nd modèle, 142 M", "22,5")):
        ax.text(0.10, y, nom, ha="left", va="center",
                color=INK_SECONDARY, fontsize=20)
        ax.text(0.90, y, val, ha="right", va="center",
                color=SERIES[0], fontsize=30, fontweight="bold")

    ax.plot([0.12, 0.88], [0.225, 0.225], color="#e6e5e1", linewidth=2)
    ax.text(0.5, 0.140, "En dessous, le modèle est sous-alimenté.", ha="center",
            color=INK_SECONDARY, fontsize=18)
    ax.text(0.5, 0.075, "Au-dessus, on paie du calcul pour peu de gain.",
            ha="center", color=INK_SECONDARY, fontsize=18)

    save(fig, os.path.join(OUT, "hero_chinchilla.png"), tight=False)


def hero_mesurer():
    """Post technique : mesurer avant d'optimiser."""
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.93, espace("J'ALLAIS PARALLÉLISER"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")
    ax.text(0.5, 0.868, espace("LA MAUVAISE ÉTAPE"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    for y, nom, val, coul in (
            (0.700, "Décompresser le fichier", "130 089", SERIES[0]),
            (0.480, "Rejouer les parties", "5 163", SERIES[1])):
        ax.text(0.5, y + 0.075, nom, ha="center",
                color=INK_PRIMARY, fontsize=22)
        ax.text(0.5, y - 0.010, val, ha="center",
                color=coul, fontsize=52, fontweight="bold")
        ax.text(0.5, y - 0.070, "parties par seconde", ha="center",
                color=INK_SECONDARY, fontsize=17)

    ax.plot([0.12, 0.88], [0.345, 0.345], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.245, "Facteur 25.", ha="center",
            color=INK_PRIMARY, fontsize=30, fontweight="bold")
    ax.text(0.5, 0.170, "Mon intuition disait l'inverse.", ha="center",
            color=INK_PRIMARY, fontsize=25)
    ax.text(0.5, 0.075, "On mesure avant d'optimiser.", ha="center",
            color=INK_SECONDARY, fontsize=20)

    save(fig, os.path.join(OUT, "hero_mesurer.png"), tight=False)


def hero_lichess():
    """Image d'ouverture du post final : la première victoire sur Lichess.

    Un seul message : un modèle qui n'a appris aucune règle a gagné une vraie
    partie contre un humain. Le score en très gros, la thèse en une ligne.
    """
    VERT = "#6a9b41"          # vert Lichess, rappel de l'endroit
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.90, espace("SUR LICHESS, CONTRE UN HUMAIN"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    ax.text(0.5, 0.70, "0 – 1", ha="center",
            color=VERT, fontsize=96, fontweight="bold")
    ax.text(0.5, 0.585, "philidor-142M gagne", ha="center",
            color=INK_SECONDARY, fontsize=22)

    ax.plot([0.14, 0.86], [0.500, 0.500], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.405, "Aucune règle apprise.", ha="center",
            color=INK_PRIMARY, fontsize=30, fontweight="bold")
    ax.text(0.5, 0.315, "Aucun coup illégal.", ha="center",
            color=INK_PRIMARY, fontsize=30, fontweight="bold")
    ax.text(0.5, 0.205, "Une fourchette de cavalier pour", ha="center",
            color=INK_SECONDARY, fontsize=21)
    ax.text(0.5, 0.150, "gagner la dame.", ha="center",
            color=INK_SECONDARY, fontsize=21)

    ax.text(0.5, 0.060, "La tactique émerge de la prédiction.", ha="center",
            color=VERT, fontsize=21, fontweight="bold")

    save(fig, os.path.join(OUT, "hero_lichess.png"), tight=False)


def hero_rating():
    """Le chiffre de la finale : le classement Rapid public, établi sur Lichess.

    Le premier Elo absolu du projet — pas un écart relatif à Stockfish sur ma
    machine, mais un classement public contre d'autres bots.
    """
    VERT = "#6a9b41"
    fig, ax = plt.subplots(figsize=TAILLE)
    cadre(ax)

    ax.text(0.5, 0.90, espace("CLASSEMENT RAPID SUR LICHESS"), ha="center",
            color=INK_MUTED, fontsize=15, fontweight="bold")

    ax.text(0.5, 0.66, "1908", ha="center",
            color=VERT, fontsize=118, fontweight="bold")
    ax.text(0.5, 0.545, "établi, contre d'autres bots", ha="center",
            color=INK_SECONDARY, fontsize=22)

    ax.plot([0.14, 0.86], [0.470, 0.470], color="#e6e5e1", linewidth=2)

    ax.text(0.5, 0.375, "Aucune règle. Aucune recherche.", ha="center",
            color=INK_PRIMARY, fontsize=27, fontweight="bold")
    ax.text(0.5, 0.290, "Juste prédire le coup suivant.", ha="center",
            color=INK_PRIMARY, fontsize=25)

    ax.text(0.5, 0.175, "9 victoires, dont une finale", ha="center",
            color=INK_SECONDARY, fontsize=20)
    ax.text(0.5, 0.120, "menée jusqu'au mat.", ha="center",
            color=INK_SECONDARY, fontsize=20)

    ax.text(0.5, 0.040, "Le premier Elo absolu du projet.", ha="center",
            color=VERT, fontsize=20, fontweight="bold")

    save(fig, os.path.join(OUT, "hero_rating.png"), tight=False)


if __name__ == "__main__":
    os.makedirs(OUT, exist_ok=True)
    print("Images d'ouverture :")
    for f in (hero_legalite, hero_qwen, hero_regles,
              hero_overfit, hero_loss_piege, hero_chinchilla,
              hero_mesurer, hero_lichess, hero_rating):
        try:
            f()
        except Exception as e:
            print(f"  ECHEC {f.__name__} : {e}")
game_gif.py139 lines

Builds an animated GIF of a game, move by move.

"""Rendu d'une partie Lichess en GIF animé + diagramme de la position clé.

Sert à illustrer la première victoire du bot philidor-142M sur Lichess
(partie 0ofezMba, geekours 0-1 philidor-142M).

    python game_gif.py --pgn "chemin.pgn" --out-dir figures

Aucune dépendance exotique : python-chess pour la logique, matplotlib pour le
rendu (pièces en glyphes Unicode pleins, colorés blanc/noir avec un liseré de
contraste), PIL pour assembler le GIF. Pas de cairosvg ni d'imageio.
"""

import argparse
import io

import chess
import chess.pgn
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib import patheffects
from PIL import Image

# Thème Lichess (vert), pour rester cohérent avec l'endroit où la partie a eu lieu.
CLAIR = "#ebecd0"
FONCE = "#779556"
SURLIGNE = "#f6f669"      # dernier coup
GLYPHE = {"k": "♚", "q": "♛", "r": "♜",
          "b": "♝", "n": "♞", "p": "♟"}


def rendu_position(board: chess.Board, dernier: chess.Move | None,
                   titre: str, sous_titre: str) -> Image.Image:
    """Rend une position en image PIL (échiquier vu côté Blancs)."""
    fig, ax = plt.subplots(figsize=(4.8, 5.3), dpi=110)
    ax.set_xlim(0, 8); ax.set_ylim(0, 8)
    ax.set_aspect("equal"); ax.axis("off")

    cases_surlignees = set()
    if dernier is not None:
        cases_surlignees = {dernier.from_square, dernier.to_square}

    for sq in chess.SQUARES:
        f, r = chess.square_file(sq), chess.square_rank(sq)
        base = CLAIR if (f + r) % 2 else FONCE
        ax.add_patch(plt.Rectangle((f, r), 1, 1, color=base))
        if sq in cases_surlignees:
            ax.add_patch(plt.Rectangle((f, r), 1, 1, color=SURLIGNE, alpha=0.55))
        piece = board.piece_at(sq)
        if piece:
            couleur = "white" if piece.color == chess.WHITE else "#111111"
            contour = "#111111" if piece.color == chess.WHITE else "#f4f4f4"
            ax.text(f + 0.5, r + 0.5, GLYPHE[piece.symbol().lower()],
                    fontsize=30, ha="center", va="center", color=couleur,
                    path_effects=[patheffects.withStroke(linewidth=2.2,
                                                         foreground=contour)])

    # coordonnées discrètes
    for f in range(8):
        ax.text(f + 0.5, -0.02, "abcdefgh"[f], ha="center", va="top",
                fontsize=8, color="#555")
    for r in range(8):
        ax.text(-0.02, r + 0.5, str(r + 1), ha="right", va="center",
                fontsize=8, color="#555")

    ax.set_title(titre, fontsize=14, fontweight="bold", pad=10, color="#222")
    fig.text(0.5, 0.045, sous_titre, ha="center", fontsize=10, color="#555")
    fig.subplots_adjust(left=0.04, right=0.98, top=0.9, bottom=0.09)

    buf = io.BytesIO()
    fig.savefig(buf, format="png")
    plt.close(fig)
    buf.seek(0)
    return Image.open(buf).convert("RGB")


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--pgn", required=True)
    p.add_argument("--out-dir", default="figures")
    p.add_argument("--gif", default="lichess_win.gif")
    p.add_argument("--cle", default="lichess_win_fork.png",
                   help="diagramme de la position clé")
    p.add_argument("--cle-ply", type=int, default=34,
                   help="demi-coup de la position clé (défaut 34 = 17...Nxg5)")
    p.add_argument("--ms-par-coup", type=int, default=850,
                   help="durée d'une image en ms (baisser pour les longues parties)")
    p.add_argument("--cle-titre", default="La position clé",
                   help="titre du diagramme de la position clé")
    args = p.parse_args()

    with open(args.pgn) as f:
        game = chess.pgn.read_game(f)
    blancs = game.headers.get("White", "?")
    noirs = game.headers.get("Black", "?")
    resultat = game.headers.get("Result", "*")
    ouverture = game.headers.get("Opening", "")

    board = game.board()
    frames = []
    # position initiale
    frames.append(rendu_position(
        board, None, f"{blancs}  vs  {noirs}",
        ouverture or "Position initiale"))

    moves = list(game.mainline_moves())
    cle_img = None
    for i, mv in enumerate(moves, start=1):
        san = board.san(mv)
        board.push(mv)
        num = (i + 1) // 2
        trait = "Blancs" if i % 2 == 1 else "Noirs"
        sous = f"{num}.{'' if i % 2 == 1 else '..'} {san}   ({trait})"
        img = rendu_position(board, mv, f"{blancs}  vs  {noirs}", sous)
        frames.append(img)
        if i == args.cle_ply:
            cle_img = rendu_position(board, mv, args.cle_titre, sous)

    # image finale : résultat
    frames.append(rendu_position(
        board, moves[-1] if moves else None,
        f"{blancs}  {resultat}  {noirs}",
        "philidor-142M gagne — 0 coup illégal"))

    out_gif = f"{args.out_dir}/{args.gif}"
    # tenir la 1re et la dernière frame plus longtemps
    durees = [1400] + [args.ms_par_coup] * (len(frames) - 2) + [4000]
    frames[0].save(out_gif, save_all=True, append_images=frames[1:],
                   duration=durees, loop=0, optimize=True)
    print(f"GIF   : {out_gif}  ({len(frames)} frames)")

    if cle_img is not None:
        out_cle = f"{args.out_dir}/{args.cle}"
        cle_img.save(out_cle)
        print(f"clé   : {out_cle}")


if __name__ == "__main__":
    main()
run2.sh124 lines

The complete pipeline of the second training run, end to end.

#!/usr/bin/env bash
# Second entraînement : modèle plus gros sur quatre mois de données.
#
#   modèle  : 20 couches, d=768, 12 têtes, MLP 2048 -> 141 589 248 paramètres
#   données : dumps Lichess 2026-04 à 2026-07 -> ~3.06 G tokens
#   ratio   : 21.6 tokens par paramètre, soit l'optimum de Chinchilla
#
# Le filtre Elo reste 1800-2600. Le relever diviserait le corpus par 2.6
# (mesuré sur 60 000 parties), or un modèle plus gros a besoin de PLUS de
# données, pas moins. Les deux pistes se contredisent frontalement.
#
# Batch 80 et non 96 : à batch 96 le pic mémoire atteint 22.8 Go sur 24, trop
# juste pour un run de vingt heures, alors que le débit est identique
# (34 598 contre 34 884 tokens/s). On ne paie donc rien pour cette marge.
#
# Usage : nohup ./run2.sh > logs/run2_pipeline.log 2>&1 &

set -u
REPO="${REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)}"
cd "$REPO"
PY=/mnt/ia-data/fine-tuning/chess/venv/bin/python
D=data
MOIS="2026-06 2026-05 2026-04"      # 2026-07 est déjà parsé
log() { echo "[$(date -u +%H:%M:%S)] $*"; }

# --- 1. Attendre les téléchargements ---------------------------------------
# On teste la TAILLE des fichiers, pas la présence d'un processus. Un pgrep
# sur un motif trouve aussi les commandes qui ne font que mentionner ce motif
# — c'est exactement ce qui a provoqué un interblocage plus tôt aujourd'hui.
log "attente des téléchargements..."
for m in $MOIS; do
  f="$D/lichess_db_standard_rated_${m}.pgn.zst"
  attendu=$(curl -sI "https://database.lichess.org/standard/lichess_db_standard_rated_${m}.pgn.zst" \
            | grep -i content-length | tr -d '\r' | awk '{print $2}')
  [ -z "$attendu" ] && { log "ERREUR : taille de $m introuvable"; exit 1; }
  while [ "$(stat -c %s "$f" 2>/dev/null || echo 0)" -lt "$attendu" ]; do sleep 30; done
  log "  $m complet ($(numfmt --to=iec "$attendu"))"
done

# --- 2. Parser les nouveaux mois -------------------------------------------
for m in $MOIS; do
  out="$D/games_uci_${m}.txt"
  if [ -s "$out" ]; then log "$m déjà parsé, on saute"; continue; fi
  log "=== parsing $m ==="
  $PY prepare_data.py parse \
    --input "$D/lichess_db_standard_rated_${m}.pgn.zst" \
    --workers 10 --target-tokens 0 --max-games 0 \
    --out "$out" --stats "logs/phase1_parse_${m}.json" \
    || { log "ECHEC parsing $m"; exit 1; }
done

# --- 3. Corpus combiné ------------------------------------------------------
log "=== fusion des quatre mois ==="
COMBINED="$D/games_uci_4mois.txt"
cat "$D/games_uci.txt" "$D/games_uci_2026-06.txt" \
    "$D/games_uci_2026-05.txt" "$D/games_uci_2026-04.txt" > "$COMBINED"
log "  $(wc -l < "$COMBINED") parties au total"

# --- 4. Encodage ------------------------------------------------------------
# Sautable comme le parsing : sur un relancement après incident, refaire huit
# minutes d'encodage pour rien serait absurde.
mkdir -p "$D/4mois"
if [ -s "$D/4mois/train.bin" ] && [ -s "$D/4mois/encode_stats.json" ]; then
  log "=== encodage déjà fait, on saute ==="
else
  log "=== encodage ==="
  $PY prepare_data.py encode \
    --games "$COMBINED" --vocab "$D/vocab.json" --outdir "$D/4mois" \
    || { log "ECHEC encodage"; exit 1; }
fi
cp -n "$D/vocab.json" "$D/4mois/vocab.json" 2>/dev/null || true

TOKENS=$($PY -c "import json;print(json.load(open('$D/4mois/encode_stats.json'))['tokens_train'])")
log "  train.bin : $TOKENS tokens"

# --- 5. Entraînement --------------------------------------------------------
# batch 80 x accumulation 2 = 40 960 tokens par step.
STEPS=$(( TOKENS / 40960 ))
log "=== entraînement : $STEPS steps pour une époque ==="

$PY train.py \
  --run-name run2 --device cuda:1 \
  --data-dir "$D/4mois" \
  --n-layer 20 --n-embd 768 --n-head 12 --mlp-hidden 2048 \
  --batch-size 80 --grad-accum 2 \
  --max-steps "$STEPS" --warmup-steps 2000 \
  --eval-interval 1000 --snapshot-every 5000 \
  || { log "ECHEC entraînement"; exit 1; }

# --- 6. Évaluation et Elo ---------------------------------------------------
# Température 0.6 et non 1.0 : le balayage fait sur run1 a montré qu'à
# température 1 le modèle échantillonne <eos> par accident et coupe les
# parties trop tôt, ce qui fausse le taux de parties complètes.
log "=== évaluation finale ==="
$PY evaluate.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab "$D/vocab.json" --val-games "$D/4mois/val_games.txt" \
  --n-legal 20000 --n-agreement 20000 --n-per-rule 500 --n-full-games 500 \
  --temperature 0.6 --max-scan 400000 \
  --out logs/eval_final_run2.json || log "ECHEC évaluation"

# Le duel qui intéresse vraiment : les deux modèles finaux face à face, avec
# assez de parties pour que la différence soit tranchée plutôt que suggérée.
# 400 parties donnent une marge d'environ ±35 Elo, contre ±90 pour 60 parties.
log "=== DUEL : run2 final contre run1 final, 400 parties ==="
$PY elo_match.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab "$D/vocab.json" --skip-stockfish \
  --vs checkpoints/run1_best.pt --vs-games 400 \
  --label-a "run2 (142 M, 4 mois)" --label-b "run1 (51 M, 1 mois)" \
  --out logs/duel_run2_vs_run1.json || log "ECHEC duel"

log "=== run2 contre les instantanés de run1 ==="
$PY elo_match.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab "$D/vocab.json" --skip-stockfish --ladder \
  --ladder-dir checkpoints --ladder-run run1 --ladder-games 100 \
  --out logs/elo_run2_vs_run1.json || log "ECHEC comparaison run1"

log "=== run2 contre Stockfish ==="
$PY elo_match.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab "$D/vocab.json" \
  --stockfish "${STOCKFISH:-$(command -v stockfish)}" \
  --levels 0 1 2 3 4 --games 200 --movetime-ms 50 \
  --out logs/elo_report_run2.json || log "ECHEC Stockfish"

log "=== RUN2 TERMINÉ ==="
post_run2_plots.sh74 lines

Canonical evaluations and charts after the second run.

#!/usr/bin/env bash
# Étape finale : évaluations canoniques de run2, puis tous les graphiques.
#
# POLITIQUE DE TEMPÉRATURE — le point le plus important de ce script.
#
# La température n'est pas un paramètre du modèle, c'est un paramètre de la
# MESURE. Deux runs comparés à des températures différentes produisent un écart
# artificiel qu'on attribue à tort au modèle. Chaque métrique a donc une
# température imposée, et une seule :
#
#   légalité par coup, accord humain, tests par règle
#       -> température 1.0. C'est le protocole de référence, celui du 97.86 %
#          publié pour run1. C'est aussi le test le plus sévère : on
#          échantillonne dans la distribution complète du modèle.
#
#   parties complètes sans coup illégal
#       -> mesuré aux DEUX températures, et les deux publiés. À 1.0 c'est le
#          protocole de référence ; à 0.6 c'est le régime d'usage réaliste.
#          L'écart entre les deux (52.5 % contre 80.5 % sur run1) est lui-même
#          un résultat, pas un détail à cacher.
#
#   matchs et Elo
#       -> température 0, jeu déterministe, comme engine.py. La diversité des
#          parties vient des coups d'ouverture aléatoires, pas du sampling.
#
# run2.sh lance son évaluation finale à 0.6 ; on la conserve comme mesure
# secondaire et on produit ici la mesure canonique à 1.0.
set -u
REPO="${REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)}"
cd "$REPO"
PY=/mnt/ia-data/fine-tuning/chess/venv/bin/python
D=data/4mois
M=/mnt/ia-data/fine-tuning/chess/logs/RUN2_DONE.marker
log() { echo "[$(date -u +%H:%M:%S)] $*"; }

log "attente de la fin de run2..."
until [ -f "$M" ]; do sleep 120; done
log "run2 terminé"

# --- Évaluation canonique, température 1.0, protocole identique à run1 ------
log "=== évaluation canonique de run2 (température 1.0) ==="
$PY evaluate.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab data/vocab.json --val-games "$D/val_games.txt" \
  --n-legal 20000 --n-agreement 20000 --n-per-rule 500 --n-full-games 500 \
  --temperature 1.0 --max-scan 400000 \
  --out logs/eval_final_run2_t10.json || log "ECHEC évaluation t=1.0"

# --- Parties complètes à 0.6, régime d'usage réaliste ----------------------
# Seule cette métrique dépend fortement de la température ; les autres sont
# calculées sur l'argmax et n'en dépendent pas. On ne relance donc que ça.
log "=== parties complètes de run2 (température 0.6) ==="
$PY evaluate.py --ckpt checkpoints/run2_best.pt --device cuda:1 \
  --vocab data/vocab.json --val-games "$D/val_games.txt" \
  --n-legal 2000 --n-agreement 2000 --n-per-rule 50 --n-full-games 500 \
  --temperature 0.6 --max-scan 100000 \
  --out logs/eval_final_run2_t06.json || log "ECHEC évaluation t=0.6"

# --- Même mesure sur run1, pour que la comparaison soit symétrique ---------
log "=== parties complètes de run1 (température 0.6) ==="
$PY evaluate.py --ckpt checkpoints/run1_best.pt --device cuda:1 \
  --vocab data/vocab.json --val-games data/val_games.txt \
  --n-legal 2000 --n-agreement 2000 --n-per-rule 50 --n-full-games 500 \
  --temperature 0.6 --max-scan 100000 \
  --out logs/eval_final_run1_t06.json || log "ECHEC évaluation run1 t=0.6"

# --- Graphiques -------------------------------------------------------------
log "=== graphiques ==="
$PY plots.py --run-name run2 --suffixe _run2 \
  --eval-glob "eval_run2_step*.json" --elo-json "elo_report_run2.json" \
  || log "ECHEC graphiques run2"
$PY plots.py --run-name run1 --comparer run1 run2 || log "ECHEC comparaisons"

log "=== TOUT EST TERMINÉ ==="
ls -la figures/
run_lichess_bot.sh124 lines

Launches the Lichess bot with the watchdog that only restarts on a real crash.

#!/usr/bin/env bash
# Watchdog : gère LE serveur d'inférence + lichess-bot, les maintient en vie, et
# garantit qu'aucun processus ne survit à un arrêt. Lancé une seule fois :
#
#   cd /chemin/vers/chess-model
#   setsid ./run_lichess_bot.sh > /mnt/ia-data/fine-tuning/chess/logs/watchdog.log 2>&1 < /dev/null &
#   disown
#
# ARCHITECTURE (voir INFERENCE_SERVER.md). Le modèle est sans état : au lieu
# d'une copie par partie, un unique serveur (infer_server.py) le charge une fois,
# et chaque partie utilise un client UCI léger (engine_client.py) qui relaie les
# coups vers ce serveur par une socket Unix. La concurrence n'est donc plus
# limitée par la mémoire.
#
# ARRÊT PROPRE : kill le PID de CE script (dans watchdog.pid). Le trap tue le
# bot, le serveur et tout worker restant, puis sort.
#   kill "$(cat /mnt/ia-data/fine-tuning/chess/logs/watchdog.pid)"
#
# POURQUOI tout ce soin sur les processus. lichess-bot lance des workers
# `multiprocessing` ; tués à moitié, ils survivent (PPID=1) et gardent le flux
# Lichess ouvert -> compte bloqué en « 429 » (un flux par compte). D'où : bot
# lancé via `setsid` (chef de session, kill -- -PGID fiable) et fonction
# nettoyer() de sécurité au démarrage ET à l'arrêt. Historique : LICHESS_BOT.md.

set -u
REPO="${REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)}"
BOT_DIR=/mnt/ia-data/fine-tuning/chess/lichess-bot
PY=/mnt/ia-data/fine-tuning/chess/venv/bin/python
LOG=/mnt/ia-data/fine-tuning/chess/logs/lichess_bot.log
SRV_LOG=/mnt/ia-data/fine-tuning/chess/logs/infer_server.log
PIDFILE=/mnt/ia-data/fine-tuning/chess/logs/watchdog.pid
SOCK=/mnt/ia-data/fine-tuning/chess/logs/infer.sock
CKPT="$REPO/checkpoints/run2_best.pt"
VOCAB="$REPO/data/vocab.json"
DEVICE=cuda:0
RESTART_DELAY=120

srv_pid=""
child=""

# Tue tout reste : serveur, bots (+ leur session), workers orphelins (PPID=1).
nettoyer() {
  for p in $(pgrep -f 'infer_server\.py' 2>/dev/null); do kill -KILL "$p" 2>/dev/null; done
  for b in $(pgrep -f 'lichess-bot\.py' 2>/dev/null); do
    local sid; sid=$(ps -o sid= -p "$b" 2>/dev/null | tr -d ' ')
    [ -n "$sid" ] && kill -TERM -- -"$sid" 2>/dev/null
  done
  sleep 1
  for b in $(pgrep -f 'lichess-bot\.py' 2>/dev/null); do
    local sid; sid=$(ps -o sid= -p "$b" 2>/dev/null | tr -d ' ')
    [ -n "$sid" ] && kill -KILL -- -"$sid" 2>/dev/null
    kill -KILL "$b" 2>/dev/null
  done
  for pid in $(ls /proc 2>/dev/null | grep -E '^[0-9]+$'); do
    local comm cmd ppid
    comm=$(cat /proc/$pid/comm 2>/dev/null) || continue
    case "$comm" in python*) ;; *) continue ;; esac
    cmd=$(tr '\0' ' ' < /proc/$pid/cmdline 2>/dev/null)
    case "$cmd" in *multiprocessing*) ;; *) continue ;; esac
    ppid=$(awk '/^PPid:/{print $2}' /proc/$pid/status 2>/dev/null)
    [ "$ppid" = "1" ] && kill -KILL "$pid" 2>/dev/null
  done
}

# (Re)démarre le serveur d'inférence et attend qu'il soit prêt (modèle chargé).
demarrer_serveur() {
  for p in $(pgrep -f 'infer_server\.py' 2>/dev/null); do kill -KILL "$p" 2>/dev/null; done
  rm -f "$SOCK" "$SOCK.ready"
  echo "[watchdog $(date -u +%H:%M:%S)] démarrage du serveur d'inférence ($DEVICE)"
  ( cd "$REPO" && setsid "$PY" infer_server.py --ckpt "$CKPT" --vocab "$VOCAB" \
      --device "$DEVICE" --socket "$SOCK" >> "$SRV_LOG" 2>&1 ) &
  srv_pid=$!
  for _ in $(seq 1 90); do          # jusqu'à ~90 s pour charger le modèle
    [ -e "$SOCK.ready" ] && { echo "[watchdog $(date -u +%H:%M:%S)] serveur prêt"; return 0; }
    sleep 1
  done
  echo "[watchdog $(date -u +%H:%M:%S)] ERREUR : serveur non prêt après 90 s" >&2
  return 1
}

serveur_vivant() {
  [ -S "$SOCK" ] && pgrep -f 'infer_server\.py' >/dev/null 2>&1
}

# Garde anti-doublon : deux watchdogs = deux flux = 429.
if [ -f "$PIDFILE" ]; then
  old=$(cat "$PIDFILE" 2>/dev/null)
  if [ -n "$old" ] && kill -0 "$old" 2>/dev/null; then
    echo "[watchdog] un watchdog tourne déjà (PID $old), abandon." >&2
    exit 1
  fi
fi

echo "[watchdog $(date -u +%H:%M:%S)] nettoyage des restes éventuels"
nettoyer
echo $$ > "$PIDFILE"

arreter() {
  echo "[watchdog $(date -u +%H:%M:%S)] arrêt demandé"
  [ -n "$child" ] && kill -- -"$child" 2>/dev/null
  [ -n "$srv_pid" ] && kill "$srv_pid" 2>/dev/null
  sleep 2
  nettoyer
  rm -f "$PIDFILE" "$SOCK" "$SOCK.ready"
  exit 0
}
trap arreter TERM INT

demarrer_serveur || { rm -f "$PIDFILE"; exit 1; }

cd "$BOT_DIR" || exit 1
while true; do
  serveur_vivant || { echo "[watchdog $(date -u +%H:%M:%S)] serveur mort, relance"; demarrer_serveur; }
  echo "[watchdog $(date -u +%H:%M:%S)] démarrage du bot"
  setsid "$PY" lichess-bot.py >> "$LOG" 2>&1 &
  child=$!                      # setsid : chef de session, PGID == $child
  wait "$child"
  code=$?
  kill -- -"$child" 2>/dev/null
  sleep 1
  # ne PAS appeler nettoyer() ici : ça tuerait le serveur ; on ne balaie que le bot
  echo "[watchdog $(date -u +%H:%M:%S)] bot terminé (code $code), redémarrage dans ${RESTART_DELAY}s"
  sleep "$RESTART_DELAY"
done
duels_qwen.sh86 lines

Chains the three protocols of the duel against the general-purpose model.

#!/usr/bin/env bash
# Enchaîne les quatre duels contre Qwen, sans intervention.
#
#   run2 x 30 tentatives   (en cours au lancement de ce script)
#   run1 x 30 tentatives   -> l'adversaire affronte des positions différentes
#   run2 x strict          -> règle réelle : un coup illégal fait perdre
#   run1 x strict
#
# Pourquoi les quatre. Le taux de légalité de Qwen dépend des positions qu'on
# lui présente, donc du modèle qu'il affronte : contre un adversaire plus
# faible, les parties partent ailleurs. Mesurer avec les deux modèles sépare ce
# qui vient de Qwen de ce qui vient de la position.
#
# Les deux runs stricts sont rapides — les parties s'arrêtent au premier coup
# illégal, soit après un ou deux coups.
#
# Usage : nohup ./duels_qwen.sh > logs/duels_qwen_pipeline.log 2>&1 &

set -u
REPO="${REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)}"
cd "$REPO"
PY=/mnt/ia-data/fine-tuning/chess/venv/bin/python
log() { echo "[$(date -u +%H:%M:%S)] $*"; }

# Attendre la fin du duel déjà lancé. Deux façons ratées avant celle-ci :
#
#   1. tester la présence du fichier de sortie -> il contenait encore le
#      rapport d'un run PRÉCÉDENT, le script a démarré aussitôt et deux duels
#      se sont disputé le serveur ;
#   2. comparer les dates du JSON et du log -> le log est écrit APRÈS le JSON,
#      donc la condition ne devenait jamais fausse.
#
# On attend le PID, passé en argument. C'est sans ambiguïté, contrairement à
# pgrep qui trouve aussi les commandes mentionnant simplement le motif — piège
# qui a déjà provoqué un interblocage sur ce projet.
PID_EN_COURS="${1:-}"
if [ -n "$PID_EN_COURS" ]; then
  log "attente du duel en cours (PID $PID_EN_COURS)..."
  while kill -0 "$PID_EN_COURS" 2>/dev/null; do sleep 60; done
  log "duel run2 (30 tentatives) terminé"
fi

verifier_serveur() {
  for _ in $(seq 1 30); do
    if curl -s -m 5 http://localhost:8080/health | grep -q '"ok"'; then return 0; fi
    log "  serveur llama indisponible, nouvelle tentative dans 30 s"
    sleep 30
  done
  log "ERREUR : serveur llama injoignable, arrêt"
  exit 1
}

lancer() {
  local ckpt=$1 nom=$2 sortie=$3; shift 3
  if [ -s "$sortie" ]; then log "$nom déjà fait, on saute"; return; fi
  verifier_serveur
  log "=== $nom ==="
  $PY -u qwen_match.py --ckpt "$ckpt" --device cuda:1 --games 100 \
      --out "$sortie" "$@" || log "ECHEC $nom"
}

lancer checkpoints/run1_best.pt "run1 contre Qwen, 30 tentatives" \
       logs/duel_qwen_run1.json

lancer checkpoints/run2_best.pt "run2 contre Qwen, RÈGLE RÉELLE" \
       logs/duel_qwen_run2_strict.json --strict

lancer checkpoints/run1_best.pt "run1 contre Qwen, RÈGLE RÉELLE" \
       logs/duel_qwen_run1_strict.json --strict

log "=== TOUS LES DUELS TERMINÉS ==="
for f in logs/duel_qwen.json logs/duel_qwen_run1.json \
         logs/duel_qwen_run2_strict.json logs/duel_qwen_run1_strict.json; do
  [ -s "$f" ] && $PY - "$f" <<'EOF'
import json, sys
d = json.load(open(sys.argv[1]))
r = d["resultat_parties"]; l = d["legalite_premiere_tentative"]
ins = d.get("insistance_necessaire", {})
print(f"\n{sys.argv[1]}")
print(f"  notre modèle : +{r['victoires']} ={r['nulles']} -{r['defaites']}"
      f"   fins : {d['terminaisons']}")
print(f"  Qwen légal au 1er essai : {l['taux']:.2%} sur {l['coups_demandes']} coups")
if ins.get("tentatives_moyennes_par_coup"):
    print(f"  tentatives moyennes : {ins['tentatives_moyennes_par_coup']}")
EOF
done
restart_llama_server.sh41 lines

Restarts the server of the general-purpose LLM used as an opponent.

#!/usr/bin/env bash
# Relance le serveur LLM de Billy exactement comme il tournait avant que ce
# projet ne réquisitionne les GPU.
#
# Ligne de commande capturée le 2026-08-04 à 13:29 UTC depuis `ps aux` sur le
# PID 1039, qui tournait alors depuis 21 h 44 min (démarré le 2026-08-03).
# Le service répondait {"status":"ok"} sur http://localhost:8080/health.
#
# Occupation VRAM constatée avant arrêt :
#   RTX 3060 (bus 06:00.0) : 8127 MiB
#   RTX 3090 (bus 07:00.0) : 17542 MiB
#
# Usage : ./restart_llama_server.sh

set -euo pipefail

LOG="${LOG:-$HOME/llama-server.log}"
LLAMA_SERVER="${LLAMA_SERVER:-$HOME/llama.cpp/build/bin/llama-server}"

nohup "$LLAMA_SERVER" \
  -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL \
  --alias qwen36-35b \
  --jinja \
  -ngl 99 \
  -fa on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  -c 131072 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0 \
  # 127.0.0.1 : le serveur n'écoute QUE en local. L'API de llama-server n'a
  # aucune authentification ; l'exposer sur 0.0.0.0 la rendrait accessible à
  # toute personne sur le réseau. À ne changer qu'en connaissance de cause.
  --host 127.0.0.1 \
  --port 8080 \
  > "$LOG" 2>&1 &

echo "llama-server relancé, PID $!, logs dans $LOG"
echo "Vérifier avec : curl -s http://localhost:8080/health"
finish.sh83 lines

The end-of-run checks: integrity, archiving, summary.

#!/usr/bin/env bash
# Enchaîne tout ce qui suit l'entraînement, sans intervention.
#
#   1. attend la fin du run
#   2. évaluation finale sur gros échantillons (le chiffre publiable)
#   3. échelle Elo des instantanés entre eux
#   4. matchs contre Stockfish, du plus faible au plus fort
#   5. regénération des graphiques
#
# Chaque étape écrit son log et son JSON. Une étape qui échoue n'empêche pas
# les suivantes : mieux vaut trois résultats sur quatre qu'un arrêt complet
# après la première erreur.
#
# Usage : nohup ./finish.sh > logs/finish.log 2>&1 &

set -u
REPO="${REPO:-$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)}"
cd "$REPO"
PY=/mnt/ia-data/fine-tuning/chess/venv/bin/python
SF="${STOCKFISH:-$(command -v stockfish)}"
CKPT=checkpoints/run1_best.pt

log() { echo "[$(date -u +%H:%M:%S)] $*"; }

log "attente de la fin de l'entraînement..."
while pgrep -f "train.py --run-name run1" > /dev/null; do sleep 30; done
log "entraînement terminé"

# On n'attend PAS que le watcher se termine de lui-même : il patiente 45
# minutes sans nouvel instantané avant de s'arrêter, ce qui retarderait tout
# d'autant. Et cette attente serait de toute façon inutile — le watcher
# travaille sur le 3060 tandis que tout ce qui suit tourne sur le 3090, donc
# les deux ne se disputent rien.
# On laisse simplement une courte grâce pour que le dernier instantané soit
# évalué et rejoigne la courbe de progression, puis on arrête le watcher.
log "grâce de 3 min pour l'évaluation du dernier instantané..."
sleep 180
if pgrep -f "eval_watcher.py" > /dev/null; then
  log "arrêt du watcher"
  pkill -f "eval_watcher.py"
fi

if [ ! -f "$CKPT" ]; then
  log "ERREUR : $CKPT introuvable, arrêt"
  exit 1
fi

# --- 2. Évaluation finale -------------------------------------------------
# Échantillons dix fois plus grands que ceux des évaluations intermédiaires :
# ici on ne trace pas une courbe, on publie un chiffre. L'intervalle de
# confiance se resserre comme la racine du nombre d'observations, donc passer
# de 2 000 à 20 000 positions divise sa largeur par un peu plus de trois.
log "=== évaluation finale ==="
$PY evaluate.py \
  --ckpt "$CKPT" --device cuda:1 \
  --n-legal 20000 --n-agreement 20000 \
  --n-per-rule 500 --n-full-games 500 \
  --max-scan 400000 \
  --out logs/eval_final.json 2>&1 || log "ECHEC de l'évaluation finale"

# --- 3. Échelle Elo des instantanés ---------------------------------------
log "=== échelle Elo des instantanés ==="
$PY elo_match.py \
  --ckpt "$CKPT" --device cuda:1 \
  --skip-stockfish --ladder --ladder-games 60 \
  --out logs/elo_ladder.json 2>&1 || log "ECHEC de l'échelle Elo"

# --- 4. Matchs contre Stockfish -------------------------------------------
# Les niveaux sont joués dans l'ordre croissant : si le temps manque, on aura
# au moins les niveaux faibles, qui sont ceux où le modèle a une chance de
# marquer des points et donc où l'Elo est mesurable plutôt que borné.
log "=== matchs contre Stockfish ==="
$PY elo_match.py \
  --ckpt "$CKPT" --device cuda:1 \
  --stockfish "$SF" \
  --levels 0 1 2 3 --games 200 --movetime-ms 50 \
  --out logs/elo_report.json 2>&1 || log "ECHEC des matchs Stockfish"

# --- 5. Graphiques ---------------------------------------------------------
log "=== graphiques ==="
$PY plots.py --run-name run1 2>&1 || log "ECHEC des graphiques"

log "=== TOUT EST TERMINÉ ==="

Glossary

Every technical term used in this article, explained without jargon. They are grouped by theme rather than alphabetically: read in order, they form a fairly complete introduction to the vocabulary of training a model.

The basics

Model. A very large set of numbers (the parameters) organised into layers of computation. You present it with an input, it produces an output. The whole challenge is finding the right values for those numbers.

Parameter. One of the model's adjustable numbers. Here, 51,397,120. It is the usual measure of a model's "size". A current commercial model has several hundred billion of them.

Token. The elementary unit the model manipulates. On text, it is generally a piece of a word. Here, I chose a token to be exactly one chess move, for example e2e4.

Vocabulary. The complete list of tokens the model knows. Mine has 1971: the 1968 geometrically possible moves, plus three special tokens.

Special tokens. Three symbols that correspond to no move, but serve to delimit and align the sequences. They occupy the first three ids of the vocabulary, which makes them easy to spot and to exclude from sampling.

bos (begin of sequence). Written <bos> in the vocabulary. Marks the start of a game. The model receives it before the first move, which gives it a starting point from which to predict, including on a fresh board where it has no history to exploit.

eos (end of sequence). Written <eos>. Marks the end of a game, and it is the most interesting of the trio: since 71% of human games end in a resignation rather than a mate, the model learns that it is often the right answer before a terminal position is even reached. So when it emits it without mate, it is faithfully imitating its data. It is also the one that, sampled by accident at a high temperature, cut my games short too early and wrongly made me believe the model could not play endgames.

pad (padding). Written <pad>. A neutral token that fills out sequences that are too short so they all have the same length within a batch, a necessary condition for processing them in parallel on the GPU. It means nothing and must never be learned as a valid prediction. In this project I avoid it almost everywhere, by grouping sequences by exact length rather than padding them: badly masked padding biases the measurements without raising a single error.

Transformer. The neural network architecture used by every current language model, ChatGPT included. Its distinctive feature is the attention mechanism, which lets it connect each element of a sequence to all the others.

Training from scratch, versus fine-tuning. Training from scratch means starting from random parameters and learning everything. Fine-tuning means taking a model already trained by someone else and specialising it. This project is a training from scratch.

Inference. Using the model once it is trained, as opposed to training it. It is what happens when you ask ChatGPT a question.

Training

Loss. The number that measures how wrong the model is. The lower it is, the better the model predicts. All of training consists of driving that number down. A useful reference value: a model answering completely at random over 1971 tokens would have a loss of ln(1971), that is 7.59.

Gradient, and gradient descent. The gradient indicates, for each parameter, which way to change it to lower the loss. Gradient descent applies those corrections in a loop. It is the learning mechanism, and there is no other.

Optimiser (AdamW). The algorithm that decides the exact size of the corrections to apply. AdamW is the current standard: it adapts the correction to each parameter according to its history.

Step. One training iteration: the model sees a batch of examples, computes its loss, and corrects its parameters once. My training run had 16,300 of them.

Batch. The number of examples processed simultaneously at each step. The larger it is, the better the GPU is used, but the more memory it consumes.

Epoch. One complete pass over the whole dataset. My model did only one epoch: it never saw the same game twice.

Overfitting. When a model memorises its training data instead of extracting general regularities from it. It becomes excellent on what it has seen and poor on everything else. Here it is deliberately provoked on 100 games, as a functional test.

Validation. A sample of data set aside, which the model never sees during training, used to check that it generalises instead of memorising.

Data leak. When information from the validation sample accidentally ends up in training. The scores become flattering and false.

Hyperparameter. A setting chosen by the human before launching (batch size, learning rate, number of layers), as opposed to the parameters, which are learned by the machine.

Lazy (lazy evaluation). Computing nothing until the result is asked for. The opposite is eager evaluation, which computes everything up front. Laziness makes it possible to work on data larger than the available memory, and to start processing without waiting for loading to finish.

Generator (Python). The language's lazy implementation: a source that produces its elements one by one, on demand, instead of building everything up front. You can walk through 92 million games without ever holding more than one at a time, so at constant memory whatever the size of the file.

Process pool. A team of independent processes working in parallel, each on one CPU core. Pool.imap_unordered hands out the tasks and returns results as soon as they arrive, without respecting the original order. Beware its appetite: it consumes its source as fast as it can, which cancels out a generator's laziness.

Checkpoint, or snapshot. A save of the model's state at a given moment. I froze one every 1000 steps, which makes it possible to reconstruct after the fact the order in which the model learned.

The architecture

Layer. One stage of computation. Information passes through the layers one after another, each transforming it a little. My model has 16.

Embedding. The translation of a token into a vector of numbers, the only form the network can manipulate. Tokens with related meaning end up with related vectors.

Tied embeddings. When the same matrix serves to translate tokens into vectors at the input and to do the reverse at the output. Economical in parameters, but it is what caused the loss anomaly at 5.97 described earlier.

Attention, and attention heads. The mechanism by which each element of a sequence goes and "looks at" the others to decide what is relevant. Several heads work in parallel, each learning a different type of relation.

Context. The number of past elements the model can consult. Here, 256 moves, far more than a complete game.

Residual connection. A shortcut that lets information cross a layer without being transformed. It is what makes it possible to train deep networks without the gradient dissolving along the way.

Normalisation (LayerNorm, RMSNorm). A rescaling of values between layers, to stop them exploding or collapsing. RMSNorm is a simplified and faster version of LayerNorm.

Pre-norm. Normalising before each sub-layer rather than after. Makes deep networks markedly easier to train.

RoPE. A way of telling the model the order of the elements, by rotating their vectors by an angle proportional to their position. Attention becomes sensitive to relative distances between two elements rather than to their absolute position.

SwiGLU. A compute block that multiplies two projections together, one acting as a tap on the other. The network thereby learns to filter what deserves to be passed on.

Logits. The raw scores the model produces for each possible token, before conversion into probabilities. Setting a logit to minus infinity makes its token impossible to choose: that is the principle of the legality mask.

BPE. The standard tokenisation of language models, which learns to merge frequent character sequences. I do not use it here, since a move already amounts to a token.

Measurement

Probability distribution. The model's real output: not a move, but a score assigned to each of the 1971 tokens, the whole adding up to 100%.

Argmax, or top-1. The token the model gives the highest probability to, in other words its best bet. Top-5 groups its five best bets.

Human agreement. The proportion of positions where the model's best bet is exactly the move a human played, measured on games never seen in training. It is a measure of imitation fidelity, not of playing strength: an engine that plays better than the humans in the corpus would get lower agreement, not higher.

Probability mass. The sum of the probabilities of a subset of tokens. The example used here: the total share the model places on legal moves, which says more than its single best bet.

Temperature. A setting applied at generation time. At 0, the model always takes its best bet. The higher it goes, the more it draws from its less probable options. It is not a parameter of the model but a parameter of the measurement, which is the nastiest source of error in this project.

Free generation. Letting the model propose any token from its vocabulary, with no constraint. It is the harshest protocol, and the one behind my headline figure.

Legality mask. Forbidding illegal moves at inference by cancelling their logits. Produces a player that never makes a mistake, but says nothing about what the model has understood.

Confidence interval (Wilson). The range within which the true value most likely sits, given the size of the sample. "97.86% [97.65 – 98.05]" means the measurement is reliable to within a few hundredths. Wilson is the method suited to percentages close to 100%.

FLOPS and TFLOPS. The number of floating-point operations per second. One TFLOPS is a thousand billion operations per second. It is the measure of a GPU's raw power.

MFU (Model FLOPs Utilization). The share of the GPU's raw power actually used by training. At 64%, nearly two thirds of the card is doing useful work. To be compared against measured power, never against the spec sheet.

bf16. A 16-bit number format instead of 32. Half the memory and much faster, for a precision sufficient for training.

VRAM. The memory on board the graphics card. It is what limits the size of the model and of the batch.

Chinchilla. A reference study on the optimal split between model size and data volume, for a given compute budget. Its rule of thumb: around 20 tokens of data per parameter.

The chess domain

UCI. Explicit notation of a move: origin square and destination square, for example g1f3. It is the format I use, because it reads without knowing the position.

SAN. The notation of chess books, contextual: Nf3 means "the knight goes to f3", without saying which one. You have to know the position to understand it, hence the cost of converting to UCI.

Elo. The rating scale for chess players. A gap of 100 points corresponds to roughly 64% of wins for the stronger player. A beginner sits around 800, a good club player around 1800. Essential point: Elo measures a gap between two players, never a strength in absolute terms. So a rating only means something relative to a reference population, such as the one on Lichess.

Anchor point. An opponent whose rating is already known, which makes it possible to convert relative gaps into an absolute rating. Without one, you know how much a model has improved, but not what it is worth on a public scale.

Stockfish. The strongest free chess engine in the world, far above the best humans. It can be throttled by levels to make a calibrated opponent.

Half-move. A move played by one side only. A 40-move game has 80 of them.


Both models are published on Hugging Face, testable in three lines of Python, and load as standard LlamaForCausalLM:

  • philidor-51m: the first, two hours of training on one month of data. 97.86% legal moves.
  • philidor-142m: the second, eighteen hours on four months of data. 98.85% legal moves, and it is the one running the Lichess bot.

The bot plays live, and you can challenge it: lichess.org/@/philidor-142M