Synthesis

Gate counts, cell mixes, and the naive vs multiply-shift trade-off

Author
Affiliation

Taeho Lee

Published

August, 2026

Other Formats

1 Synthesis flows

Three flows are measured, all with Yosys 0.65:

Flow Command Meaning
Generic proc; synth; stat -json ev-compatible schema, num_cells maps to SynthesisMetrics.gate_count
Gate-level abc -g AND,NAND,OR,NOR,XOR,XNOR; stat 2-input gate library estimate
FPGA synth_ice40; stat iCE40 LUT4 and carry cells

The generic flow mirrors the canonical invocation in the ev verification CLI (ev/src/synth/backends/yosys.rs), so the numbers are directly comparable to other synthesis results in the SSCCS stack.

2 Gate counts

The gate counts by flow are shown in Figure Figure 1, with the ~300 gate claim region shaded.

Figure 1: Gate count by synthesis flow: shift-subtract (naive) vs multiply-shift. The shaded region is below the ~300 gate claim; the ratios mark the cost of timing closure.
Metric shift-subtract (naive) multiply-shift (current)
Generic cells (ev schema) 206 478
Gate-level estimate (2-input) 232 588
iCE40 LUT4 + SB_CARRY 95 + 66 201 + 41
Registers 0 0

Cell mix for the multiply-shift version:

Flow Mix
Generic (478) ANDNOT 25, AND 82, MUX 15, NAND 135, NOR 9, NOT 6, ORNOT 15, OR 44, XNOR 76, XOR 71
Gate-level (588) AND 113, NAND 192, NOR 29, NOT 23, OR 44, XNOR 84, XOR 103

2.1 The primitive’s other units

The same 2-input gate library estimates the two smaller units that complete the primitive:

Unit Cells (2-input) Structure
compose 258 axes to index, two constant multiplies (21i, 28p)
distance 1329 two decoder instances and three subtractors

The distance figure carries two decoders, because a pair needs both operands decoded; the comparators and subtractors around them are the remaining about 153 cells.

3 The trade-off

The ~300 gate claim in the Tagma whitepaper holds for the naive shift-subtract decoder (206 to 232 cells, below 300). It does not hold for the version that closes timing on a 12 MHz board: the multiply-shift structure costs 478 to 588 cells. The optimization is a deliberate exchange of about 2.5x gates for timing closure. The clock closure is shown in Figure Figure 2.

Figure 2: FPGA clock frequency before and after the multiply-shift optimization. The shaded region is above the 12 MHz board clock; the margin is the headroom of the optimized version.
Metric shift-subtract (naive) multiply-shift (current)
Logic levels 72 33
Critical path (UP5K) 115.90 ns 59.55 ns
Fmax 8.63 MHz 16.79 MHz
Meets 12 MHz board clock no yes

The one-cycle latency holds for both structures; what differs is the achievable clock rate.

4 Standard cell (Sky130)

The counts above come from the generic Yosys library and the iCE40 flow, not from a PDK standard cell library. hw/openroad/run.sh closes that gap: it runs ORFS end to end (synth, floorplan, place, cts, route, finish) against the SkyWater 130nm HD library, and reproduces on x86 and on Apple Silicon under Rosetta. Measured for the registered demo top at the 12 MHz board-equivalent clock (83.33 ns):

Metric Value
design area 4631 um^2, 35% utilization
critical path delay 11.19 ns
worst setup slack +72.33 ns
total power 0.897 mW (combinational 95.5%, sequential 3.3%, clock 1.2%)
pure decoder 388 cells, 2826 um^2 (yosys stat -liberty)

The 388-cell pure decoder count is the abc mapping of the multiply-shift network against sky130_fd_sc_hd__tt_025C_1v80.lib, consistent with the 478 generic cells and the 588 gate-level estimate above. The hw CI job runs the same flow and uploads the reports as the sky130-reports artifact.

Two flow bugs surfaced during setup, recorded in the derisking devlog. The clock was first declared as a virtual clock with no source port, so CTS built no clock tree; the fix gives the clock its source port. The bundled kepler-formal LEC step crashes with an illegal instruction on the available CPUs, so it is disabled with LEC_CHECK = 0; the RTL-to-netlist equivalence it would check is already proven by the yosys equiv gate.