Results

The model can pass. It is not yet good.

We train a model to design a small hardware block, and we grade every attempt with chip tools. After training, a 27-billion-parameter model passed tasks the same model could not pass before. On the tasks it passed, its circuits were still much larger than a strong coding agent’s. These numbers are from September 2026.

Two pieces, one loop

The service is a loop with two pieces. The harness is the judge. An agent writes SystemVerilog, then simulators, lint, and a place-and-route flow score the circuit on correctness, area, clock, and numeric error. A failed check scores zero. The training is the memory. Graded runs become examples, and a model is trained on them so the next agent starts closer to a design that passes.

A third piece sits on top of the harness: a written working method. It tells the agent to read the score formula first, keep a design that already passes, and test several ideas at once, because a mental model of synthesis is often wrong by a factor of two. Frontier models that follow this method set the high scores below. The trained model was tested without it.

Training: a math unit the tools can exhaust

The training task is a small arithmetic block used in softmax: an exponential, in bfloat16 and a few other formats. Every input can be checked. The tools are Verilator, Yosys, and nextpnr for a Lattice ECP5. Area and clock in the tables are what that software flow reports after place and route. They are not measurements from a board.

We fine-tuned Qwen 27B on 132 graded traces, about 5 million supervised tokens, for two epochs. The checkpoint of record is a full-weight update, not a small adapter. At test time the model worked alone for 60 minutes per run, with no written method and no coach. Sixteen variants, two runs each.

Pass rates on the exponential-unit suite, September 2026.
ModelWhat happened
Qwen 27B, no training0 of 10 runs passed. None of them submitted a design. Five tasks, two runs each.
Qwen 27B after training12 of 32 runs passed every check. 25 of 32 were numerically correct. 9 of 16 variants had at least one pass.
A variant held out of trainingBoth runs passed. Tight lookup-table limit, 140 MHz target, block memory priced differently from the training tasks.
9B models, same kind of data0 or 1 pass out of 32 runs. Size was the step that made passes possible.

When a trained run failed, the arithmetic was often already right. Of 20 failed runs, 13 were functionally correct and missed a performance check, usually the clock. A multiplier that does not fit a short pipeline shows up here again and again.

Three variants, so the gap has a shape. Score is the harness score; a missed check is zero. Astra is a strong coding agent, 50 minutes, also without the written method.
VariantTrained 27BAstra, 50 min
8-bit input, 200 MHz targetPassed. Best score 38.8, about 80% of Astra.48.6
Held-out, tight area, 140 MHzPassed both runs. Best score 12.4, about 67% of Astra.18.4
Base bfloat16 exp, 80 MHz, exactNumerically correct. Clock was 52 MHz, under the gate, so the score is 0.21.0, and it passed.

Across the nine variants the student passed, its best score was about a quarter of Astra’s, and about a fifth of the best run that used the written method. Passing designs were often three to ten times larger. The model has learned to produce a circuit the tools accept. It has not learned to produce a small one.

The method, on the same tasks

The high scores come from a frontier model following the written method, searching by measurement. On the base exponential, that run scored 28.9. Astra alone, in 50 minutes, scored 21.0. The trained student’s best correct-but-slow attempt on that task scored 2.4, and it did not pass the clock. The method is part of the product because the model, by itself, does not yet close that gap.

We have not shown that the method improves itself from one generation to the next. What we have shown is the first half of that idea: tool-graded traces in, and a model that can pass tasks the base model could not.

The harness, on a whole processor

The same style of judge also grades larger designs. On ChipBench, a public suite, frontier coding agents were asked to build a 32-bit RISC-V processor from a spec: RV32IM, with hazards, traps, and interrupts. One run per agent. Public tests only. A hidden test suite does not exist yet. These are not our trained weights.

RV32IM builds, 14 September 2026. The grader copied in only the agent’s RTL.
Claude Fable 5.1GPT-6 Astra
VerdictPassed, including three formal proofsPassed, including three formal proofs
Time to a sealed design1 hour 9 minutes2 hours 3 minutes
CoreMark IPC0.8370.771
Lookup tables, after routing6,0078,911
Estimated clock44.8 MHz27.6 MHz

A smaller core, fourteen instructions, showed the same pattern in an earlier pilot: the agent that overlapped instruction fetch with data access also used the fewest lookup tables. Protocol checks rejected a design whose instruction results looked right. That is why the harness grades the interface and the silicon cost, not only the architectural answer.

What we have not done

The point of a pilot is to replay one of your finished blocks inside this loop, graded by your checks. Where that work is allowed to run is a separate page.