Hardware

Verilog Hardware Implementation of the Tagma Coordinate Space

Author
Affiliation

SSCCS Initiative

Other Formats

1 Reference Implementation

The hardware realization of the Tagma primitive lives under hw/ in the syntagma repository. The toolchain is fully open source: Verilator for simulation, Yosys for synthesis, nextpnr-ice40 for place and route, icepack for the bitstream, icetime for timing analysis, and OpenRAM for the memory macro. The demo target is the Upduino 3.1 (iCE40UP5K-SG48).

The public surface of the hardware reference:

Artifact Purpose Location
tagma_decoder 3-axis combinational decoder hw/rtl/tagma_decoder.v
tagma_decoder_tb Exhaustive testbench, two modes hw/rtl/tagma_decoder_tb.v
tagma_compose Axis-to-index compose, the decoder inverse hw/rtl/tagma_compose.v
tagma_dist Field-wise distance of two code points hw/rtl/tagma_dist.v
tagma_segment_store Segment store behavioral model, 11,172 x 16-bit hw/rtl/tagma_segment_store.v
tagma_demo_top Board demo, registered outputs hw/rtl/tagma_demo_top.v
golden_anchors.hex 11,172 vectors from the Rust reference generated by make golden-export
check_golden_anchors.py Consistency gate over the anchor file hw/tools/
synthesis flows generic, gate-level, iCE40 hw/synth/yosys/synth_*.ys
PnR flow bitstream and timing report hw/synth/yosys/run_pnr.sh
chton_sram.py OpenRAM configuration hw/openram/

The reference definition is the Rust tagma_core; the hardware is checked against it on every input, not assumed equal.

2 Decoder RTL: the 16-bit Hardware Contract

The decoder is a pure combinational function over one 16-bit input. Given a code point in [U+AC00, U+D7A3], it produces the three structural axes. The implementation uses multiply-shift constant division, exact over the valid domain:

// hw/rtl/tagma_decoder.v
wire [15:0] sub = code - 16'hAC00;
wire [13:0] offset = sub[13:0];

// Split offset = 4*q4 + r2; division by 28 becomes division by 7 on q4.
wire [11:0] q4 = offset[13:2];
wire [1:0]  r2 = offset[1:0];

// j = offset / 28 in [0, 398]: (q4 * 9363) >> 16.
wire [24:0] jp = {13'd0, q4} * 25'd9363;
wire [8:0]  j  = jp[24:16];

// ii = offset / 588 = j / 21 in [0, 18]: (j * 781) >> 14.
wire [18:0] ip = {10'd0, j} * 19'd781;
wire [4:0]  ii = ip[18:14];

assign i = ii;
assign m = ({1'd0, j} - {5'd0, ii} * 10'd21)[4:0];
assign f = (q4 - {3'd0, j} * 12'd7)[1:0] * 4 + r2;

There are no registers, so the latency is exactly one cycle by construction. Inputs below U+AC00 underflow and are outside the valid domain; inputs in U+D7A4..U+D7AF are structurally invalid filler positions.

2.1 Why 16 Bits?

The 16-bit width is the same contract as the software Coord: a single machine register, a single code unit in UTF-16, a single SRAM word. The hardware decoder loads, decodes, and validates in one cycle. The enabling conditions are documented in the hardware index: BMP membership, the composition algorithm, contiguity of the block, and the Unicode stability policy. Together they make the whole address space fit one u16.

3 Why Multiply-Shift?

The naive implementation of the constant divisions uses shift-subtract dividers (what Yosys techmaps from / and %). It matches the ~300 gate claim of the whitepaper but does not close timing on a 12 MHz board:

Structure Depth Critical path Fmax Gates (generic)
shift-subtract (naive) 72 levels 115.90 ns 8.63 MHz 206
multiply-shift (current) 33 levels 59.55 ns 16.79 MHz 478

The multiply-shift constants are exact over the valid domain: division by 28 becomes division by 7 on q4 = offset >> 2 (q4 at most 2792), and division by 21 applies to j at most 398. Exhaustive simulation and formal equivalence prove the two networks functionally identical over all 2^16 inputs; the choice is purely a gates-versus-timing exchange.

4 Verification Pipeline

The pipeline has a single source of truth: tagma_core in Rust. The exporter writes the full coordinate decomposition into golden_anchors.hex, one packed 29-bit value per line:

offset[28:15] initial[14:10] medial[9:5] final[4:0]

Four independent channels consume it:

Channel Command Coverage
Formula simulation make sim all 11,172 code points
Golden anchors make sim-golden all 11,172, against the Rust reference
Gate-level netlist make gatesim all 11,172, against the same anchors
Formal equivalence make equiv all 2^16 inputs, RTL vs netlist
Compose channels make sim-compose, sim-compose-golden, gatesim-compose all 32^3 axis combinations, 11,172 anchors
Distance channels make sim-dist, sim-dist-golden, gatesim-dist 11,172 code points against two references and a varying-operand sweep
Compose and distance equivalence make equiv 2^15 axis combinations, 2^32 input pairs

tools/check_golden_anchors.py gates the generated file against the decomposition contract, so a stale or corrupted regeneration fails the pipeline before any simulation runs.

5 FPGA Demo Reference

The demo top registers the decode on the board clock and drives the onboard green LED through an active-low signal:

// hw/rtl/tagma_demo_top.v
always @(posedge clk) begin
    i     <= i_c;
    m     <= m_c;
    f     <= f_c;
    valid <= (code >= 16'hAC00) && (code <= 16'hD7A3);
end
assign led_valid = ~valid;   // negative-logic LED, lights on valid

Pin constraints map the 16-bit code point to GPIO switches, the axes to LED groups, and led_valid to the onboard green LED (pin 39), verified against the Upduino v3.1 reference PCF.

6 Memory: chton SRAM Reference

The coordinate space maps onto physical memory as a flat array, addressed by the 14-bit offset from the decoder. The chton segment store is a single-port SRAM macro configured for SkyWater 130nm, and its macro generation is pending an OpenRAM PDK install. What is verified today is a behavioral model of the same address map, checked exhaustively over all 11,172 slots with the reserved addresses reading zero. The OpenRAM configuration is:

# hw/openram/chton_sram.py
word_size = 16
num_words = 11172       # full Tagma space: U+AC00..U+D7A3
num_rw_ports = 1
tech_name = "sky130"

11,172 is not a power of two; the fallback is 16,384 words (2^14) with 5,212 reserved entries, since the 14-bit address space covers the valid range either way.

7 Key Engineering Decisions

Decision Rationale Consequence
Multiply-shift over shift-subtract 72 to 33 logic levels, closes 12 MHz 206 to 478 gates (generic schema)
Golden anchors from Rust, regenerated single source of truth, no drift file is a build artifact, CI regenerates
Four verification channels formula, reference, netlist, equivalence RTL and netlist proven identical on 2^16 inputs
Registered demo outputs one-cycle latency measurable with a real clock Fmax and setup checkable via icetime
Active-low LED via ~valid onboard LED is negative logic separate led_valid port, no semantic flip
Explicit width-matched Verilog Verilator -Wall clean without hiding truncation concat-based zero-extension, lint only for unused bits
ev-compatible stat -json flow numbers comparable across the SSCCS stack can be emitted as neXus Facts later

8 Quality Metrics

The reference is verified through 11,172 exhaustive vectors in three simulation channels plus formal equivalence over 2^16 inputs, with a Python consistency gate on the generated anchors. The compose and distance units run the same shape of checks over their own input spaces: formula and golden simulation, gate-level netlist simulation, and formal equivalence, over all 32^3 axis combinations and all 2^32 distance input pairs. The benchmark center sw/rust/benches/bench_hw.rs documents the software baseline in comments (single decode 1.44 ns, full space 1.92 µs, golden export 4.56 µs). The whole gate runs with make -C hw check and is wired into run.sh.

Measured hardware figures (branch 48-hw-verification):

Metric Value
generic cells (ev schema) 478, 0 registers
gate-level estimate 588 (2-input library)
iCE40 201 LUT4 + 41 SB_CARRY
PnR ICESTORM_LC 255 / 5280 (4%)
critical path 59.55 ns
Fmax 16.79 MHz, closes 12 MHz board clock
logic levels 33

9 Development Experience

The hardware track produced findings that changed the framing of the primitive:

  • The exhaustive test exposed that the last valid syllable is U+D7A3, not U+D7AF. Every document stating the range as ending at 0xD7AF should be checked; the tagma_core reference documents the filler positions U+D7A4..U+D7AF.
  • The 16-bit contract rests on four independent conditions (BMP membership, composition algorithm, contiguity, stability policy). BMP alone is necessary but not sufficient.
  • Timing closure costs about 2.5x gates: the ~300 gate claim holds for the naive structure, and the demo version exceeds it. The one-cycle latency holds for both.
  • The hardware is not a speedup for the N=1 dense case. The whole 11,172-entry space fits in L1 cache, which the software already exploits at 0.38 ns per access. The hardware value is energy, determinism, and physical embedding, not per-operation speed.
  • Toolchain notes: Verilator requires explicit width-matched operands and lint scoping for intentional truncation; icetime needs the iCE40 chipdb discoverable (a Homebrew symlink on macOS); the golden anchor file is a language-neutral contract that can gate a C++ port later.

10 Next Steps

Track Step Dependency
Phase 3 Physical Upduino 3.1 bring-up: program the bitstream, wire DIP switch and LED modules per the PCF, capture the demo video board and modules
Phase 4 OpenROAD standard cell flow on Sky130: timing, area, power report. The flow is set up and runs end to end (synth, floorplan, place, cts, route, finish), measured locally and re-run in the syntagma hw CI job. Two flow bugs were fixed during derisking (virtual clock, kepler-formal LEC crash), see the devlog entry none
chton propagation Generate the segment store SRAM from chton_sram.py (OpenRAM) and attach it to the decoder OpenRAM and PDK
Power model Run the OpenROAD power analysis on the VCD trace from make sim-trace Phase 4
Multi-target Port the demo flow to ECP5 or Artix-7 board choice
Streaming Pipelined decode at one code point per clock; the acceleration case appears at N greater than 1 follow-up design

11 Validation and Status

Verified: the decoder matches the Rust reference on all 11,172 valid syllables in three independent channels, the RTL and the synthesized netlist are formally equivalent over all 2^16 inputs, and the demo closes the 12 MHz board clock with margin (16.79 MHz measured). The compose and distance units are verified the same way over their own input spaces. The segment store behavioral model is verified over all 11,172 slots. The bitstream is generated and ready for board bring-up.

Pending: physical board demonstration and video, the CI confirmation of the OpenROAD standard cell report (measured locally: 4631 um^2, 11.19 ns critical path, 0.897 mW for the demo top; 388 Sky130 cells, 2826 um^2 for the pure decoder), the OpenRAM SRAM macro, and the power model on the VCD activity trace. These are follow-up tracks, not blockers for the current reference.

12 References

The companion whitepapers describe the primitive definition and the software reference implementation; the hardware index carries the measured results summary.