Results
The model can pass. It is not yet good.
We train a model to design a small hardware block, and we grade every attempt with chip tools. After training, a 27-billion-parameter model passed tasks the same model could not pass before. On the tasks it passed, its circuits were still much larger than a strong coding agent’s. These numbers are from September 2026.
Two pieces, one loop
The service is a loop with two pieces. The harness is the judge. An agent writes SystemVerilog, then simulators, lint, and a place-and-route flow score the circuit on correctness, area, clock, and numeric error. A failed check scores zero. The training is the memory. Graded runs become examples, and a model is trained on them so the next agent starts closer to a design that passes.
A third piece sits on top of the harness: a written working method. It tells the agent to read the score formula first, keep a design that already passes, and test several ideas at once, because a mental model of synthesis is often wrong by a factor of two. Frontier models that follow this method set the high scores below. The trained model was tested without it.
Training: a math unit the tools can exhaust
The training task is a small arithmetic block used in softmax: an exponential, in bfloat16 and a few other formats. Every input can be checked. The tools are Verilator, Yosys, and nextpnr for a Lattice ECP5. Area and clock in the tables are what that software flow reports after place and route. They are not measurements from a board.
We fine-tuned Qwen 27B on 132 graded traces, about 5 million supervised tokens, for two epochs. The checkpoint of record is a full-weight update, not a small adapter. At test time the model worked alone for 60 minutes per run, with no written method and no coach. Sixteen variants, two runs each.
| Model | What happened |
|---|---|
| Qwen 27B, no training | 0 of 10 runs passed. None of them submitted a design. Five tasks, two runs each. |
| Qwen 27B after training | 12 of 32 runs passed every check. 25 of 32 were numerically correct. 9 of 16 variants had at least one pass. |
| A variant held out of training | Both runs passed. Tight lookup-table limit, 140 MHz target, block memory priced differently from the training tasks. |
| 9B models, same kind of data | 0 or 1 pass out of 32 runs. Size was the step that made passes possible. |
When a trained run failed, the arithmetic was often already right. Of 20 failed runs, 13 were functionally correct and missed a performance check, usually the clock. A multiplier that does not fit a short pipeline shows up here again and again.
| Variant | Trained 27B | Astra, 50 min |
|---|---|---|
| 8-bit input, 200 MHz target | Passed. Best score 38.8, about 80% of Astra. | 48.6 |
| Held-out, tight area, 140 MHz | Passed both runs. Best score 12.4, about 67% of Astra. | 18.4 |
| Base bfloat16 exp, 80 MHz, exact | Numerically correct. Clock was 52 MHz, under the gate, so the score is 0. | 21.0, and it passed. |
Across the nine variants the student passed, its best score was about a quarter of Astra’s, and about a fifth of the best run that used the written method. Passing designs were often three to ten times larger. The model has learned to produce a circuit the tools accept. It has not learned to produce a small one.
The method, on the same tasks
The high scores come from a frontier model following the written method, searching by measurement. On the base exponential, that run scored 28.9. Astra alone, in 50 minutes, scored 21.0. The trained student’s best correct-but-slow attempt on that task scored 2.4, and it did not pass the clock. The method is part of the product because the model, by itself, does not yet close that gap.
We have not shown that the method improves itself from one generation to the next. What we have shown is the first half of that idea: tool-graded traces in, and a model that can pass tasks the base model could not.
The harness, on a whole processor
The same style of judge also grades larger designs. On ChipBench, a public suite, frontier coding agents were asked to build a 32-bit RISC-V processor from a spec: RV32IM, with hazards, traps, and interrupts. One run per agent. Public tests only. A hidden test suite does not exist yet. These are not our trained weights.
| Claude Fable 5.1 | GPT-6 Astra | |
|---|---|---|
| Verdict | Passed, including three formal proofs | Passed, including three formal proofs |
| Time to a sealed design | 1 hour 9 minutes | 2 hours 3 minutes |
| CoreMark IPC | 0.837 | 0.771 |
| Lookup tables, after routing | 6,007 | 8,911 |
| Estimated clock | 44.8 MHz | 27.6 MHz |
A smaller core, fourteen instructions, showed the same pattern in an earlier pilot: the agent that overlapped instruction fetch with data access also used the fewest lookup tables. Protocol checks rejected a design whose instruction results looked right. That is why the harness grades the interface and the silicon cost, not only the architectural answer.
What we have not done
- We have not programmed an FPGA board we own. Clock and area come from Yosys and nextpnr, a software model of a Lattice ECP5. A Linux boot on a physical board is not a result.
- We have not run this loop on a customer’s process, cell library, or sign-off tools. The FPGA flow is not an ASIC flow. Place-and-route seeds are not process corners.
- Each CPU number is one run. The training eval is two runs per variant. That is enough to see a gap. It is not a leaderboard.
The point of a pilot is to replay one of your finished blocks inside this loop, graded by your checks. Where that work is allowed to run is a separate page.
