Measurements

What Training Changed, and What It Did Not

A fine-tuned 27B model passed hardware tasks the base model could not. Its circuits are still much larger than a strong agent's, and none of this was measured on a physical board.

The task is a small exponential unit, the kind used in softmax, graded on every input. Simulators check the numbers. Yosys and nextpnr, aimed at a Lattice ECP5, report area and clock. That is a software model of an FPGA. We have not loaded a bitstream onto a board we own.

Qwen 27B with no training passed 0 of 10 runs and never submitted a design. After a full-weight fine-tune on 132 tool-graded traces, the same model passed 12 of 32 hour-long runs, covering 9 of 16 variants. A variant left out of training passed both of its runs. The same recipe at 9B parameters almost never passed.

The misses are mostly performance, not arithmetic. Thirteen of twenty failed runs were numerically correct and missed a gate, usually the clock. Across the variants it did pass, its best score was about a quarter of a strong coding agent's, and the circuits were often three to ten times larger. The full tables, and the written method that still beats the trained model, are on the results page.

All notes