Other Formats
Hardware
Verilog Hardware Implementation of the Tagma Coordinate Space
1 Reference Implementation
The hardware realization of the Tagma primitive lives under hw/ in the syntagma repository. The toolchain is fully open source: Verilator for simulation, Yosys for synthesis, nextpnr-ice40 for place and route, icepack for the bitstream, icetime for timing analysis, and OpenRAM for the memory macro. The demo target is the Upduino 3.1 (iCE40UP5K-SG48).
The public surface of the hardware reference:
| Artifact | Purpose | Location |
|---|---|---|
tagma_decoder |
3-axis combinational decoder | hw/rtl/tagma_decoder.v |
tagma_decoder_tb |
Exhaustive testbench, two modes | hw/rtl/tagma_decoder_tb.v |
tagma_compose |
Axis-to-index compose, the decoder inverse | hw/rtl/tagma_compose.v |
tagma_dist |
Field-wise distance of two code points | hw/rtl/tagma_dist.v |
tagma_segment_store |
Segment store behavioral model, 11,172 x 16-bit | hw/rtl/tagma_segment_store.v |
tagma_demo_top |
Board demo, registered outputs | hw/rtl/tagma_demo_top.v |
golden_anchors.hex |
11,172 vectors from the Rust reference | generated by make golden-export |
check_golden_anchors.py |
Consistency gate over the anchor file | hw/tools/ |
| synthesis flows | generic, gate-level, iCE40 | hw/synth/yosys/synth_*.ys |
| PnR flow | bitstream and timing report | hw/synth/yosys/run_pnr.sh |
chton_sram.py |
OpenRAM configuration | hw/openram/ |
The reference definition is the Rust tagma_core; the hardware is checked against it on every input, not assumed equal.
2 Decoder RTL: the 16-bit Hardware Contract
The decoder is a pure combinational function over one 16-bit input. Given a code point in [U+AC00, U+D7A3], it produces the three structural axes. The implementation uses multiply-shift constant division, exact over the valid domain:
// hw/rtl/tagma_decoder.v
wire [15:0] sub = code - 16'hAC00;
wire [13:0] offset = sub[13:0];
// Split offset = 4*q4 + r2; division by 28 becomes division by 7 on q4.
wire [11:0] q4 = offset[13:2];
wire [1:0] r2 = offset[1:0];
// j = offset / 28 in [0, 398]: (q4 * 9363) >> 16.
wire [24:0] jp = {13'd0, q4} * 25'd9363;
wire [8:0] j = jp[24:16];
// ii = offset / 588 = j / 21 in [0, 18]: (j * 781) >> 14.
wire [18:0] ip = {10'd0, j} * 19'd781;
wire [4:0] ii = ip[18:14];
assign i = ii;
assign m = ({1'd0, j} - {5'd0, ii} * 10'd21)[4:0];
assign f = (q4 - {3'd0, j} * 12'd7)[1:0] * 4 + r2;There are no registers, so the latency is exactly one cycle by construction. Inputs below U+AC00 underflow and are outside the valid domain; inputs in U+D7A4..U+D7AF are structurally invalid filler positions.
2.1 Why 16 Bits?
The 16-bit width is the same contract as the software Coord: a single machine register, a single code unit in UTF-16, a single SRAM word. The hardware decoder loads, decodes, and validates in one cycle. The enabling conditions are documented in the hardware index: BMP membership, the composition algorithm, contiguity of the block, and the Unicode stability policy. Together they make the whole address space fit one u16.
3 Why Multiply-Shift?
The naive implementation of the constant divisions uses shift-subtract dividers (what Yosys techmaps from / and %). It matches the ~300 gate claim of the whitepaper but does not close timing on a 12 MHz board:
| Structure | Depth | Critical path | Fmax | Gates (generic) |
|---|---|---|---|---|
| shift-subtract (naive) | 72 levels | 115.90 ns | 8.63 MHz | 206 |
| multiply-shift (current) | 33 levels | 59.55 ns | 16.79 MHz | 478 |
The multiply-shift constants are exact over the valid domain: division by 28 becomes division by 7 on q4 = offset >> 2 (q4 at most 2792), and division by 21 applies to j at most 398. Exhaustive simulation and formal equivalence prove the two networks functionally identical over all 2^16 inputs; the choice is purely a gates-versus-timing exchange.
4 Verification Pipeline
The pipeline has a single source of truth: tagma_core in Rust. The exporter writes the full coordinate decomposition into golden_anchors.hex, one packed 29-bit value per line:
offset[28:15] initial[14:10] medial[9:5] final[4:0]
Four independent channels consume it:
| Channel | Command | Coverage |
|---|---|---|
| Formula simulation | make sim |
all 11,172 code points |
| Golden anchors | make sim-golden |
all 11,172, against the Rust reference |
| Gate-level netlist | make gatesim |
all 11,172, against the same anchors |
| Formal equivalence | make equiv |
all 2^16 inputs, RTL vs netlist |
| Compose channels | make sim-compose, sim-compose-golden, gatesim-compose |
all 32^3 axis combinations, 11,172 anchors |
| Distance channels | make sim-dist, sim-dist-golden, gatesim-dist |
11,172 code points against two references and a varying-operand sweep |
| Compose and distance equivalence | make equiv |
2^15 axis combinations, 2^32 input pairs |
tools/check_golden_anchors.py gates the generated file against the decomposition contract, so a stale or corrupted regeneration fails the pipeline before any simulation runs.
5 FPGA Demo Reference
The demo top registers the decode on the board clock and drives the onboard green LED through an active-low signal:
// hw/rtl/tagma_demo_top.v
always @(posedge clk) begin
i <= i_c;
m <= m_c;
f <= f_c;
valid <= (code >= 16'hAC00) && (code <= 16'hD7A3);
end
assign led_valid = ~valid; // negative-logic LED, lights on validPin constraints map the 16-bit code point to GPIO switches, the axes to LED groups, and led_valid to the onboard green LED (pin 39), verified against the Upduino v3.1 reference PCF.
6 Memory: chton SRAM Reference
The coordinate space maps onto physical memory as a flat array, addressed by the 14-bit offset from the decoder. The chton segment store is a single-port SRAM macro configured for SkyWater 130nm, and its macro generation is pending an OpenRAM PDK install. What is verified today is a behavioral model of the same address map, checked exhaustively over all 11,172 slots with the reserved addresses reading zero. The OpenRAM configuration is:
# hw/openram/chton_sram.py
word_size = 16
num_words = 11172 # full Tagma space: U+AC00..U+D7A3
num_rw_ports = 1
tech_name = "sky130"11,172 is not a power of two; the fallback is 16,384 words (2^14) with 5,212 reserved entries, since the 14-bit address space covers the valid range either way.
7 Key Engineering Decisions
| Decision | Rationale | Consequence |
|---|---|---|
| Multiply-shift over shift-subtract | 72 to 33 logic levels, closes 12 MHz | 206 to 478 gates (generic schema) |
| Golden anchors from Rust, regenerated | single source of truth, no drift | file is a build artifact, CI regenerates |
| Four verification channels | formula, reference, netlist, equivalence | RTL and netlist proven identical on 2^16 inputs |
| Registered demo outputs | one-cycle latency measurable with a real clock | Fmax and setup checkable via icetime |
Active-low LED via ~valid |
onboard LED is negative logic | separate led_valid port, no semantic flip |
| Explicit width-matched Verilog | Verilator -Wall clean without hiding truncation |
concat-based zero-extension, lint only for unused bits |
ev-compatible stat -json flow |
numbers comparable across the SSCCS stack | can be emitted as neXus Facts later |
8 Quality Metrics
The reference is verified through 11,172 exhaustive vectors in three simulation channels plus formal equivalence over 2^16 inputs, with a Python consistency gate on the generated anchors. The compose and distance units run the same shape of checks over their own input spaces: formula and golden simulation, gate-level netlist simulation, and formal equivalence, over all 32^3 axis combinations and all 2^32 distance input pairs. The benchmark center sw/rust/benches/bench_hw.rs documents the software baseline in comments (single decode 1.44 ns, full space 1.92 µs, golden export 4.56 µs). The whole gate runs with make -C hw check and is wired into run.sh.
Measured hardware figures (branch 48-hw-verification):
| Metric | Value |
|---|---|
| generic cells (ev schema) | 478, 0 registers |
| gate-level estimate | 588 (2-input library) |
| iCE40 | 201 LUT4 + 41 SB_CARRY |
| PnR ICESTORM_LC | 255 / 5280 (4%) |
| critical path | 59.55 ns |
| Fmax | 16.79 MHz, closes 12 MHz board clock |
| logic levels | 33 |
9 Development Experience
The hardware track produced findings that changed the framing of the primitive:
- The exhaustive test exposed that the last valid syllable is U+D7A3, not U+D7AF. Every document stating the range as ending at 0xD7AF should be checked; the
tagma_corereference documents the filler positions U+D7A4..U+D7AF. - The 16-bit contract rests on four independent conditions (BMP membership, composition algorithm, contiguity, stability policy). BMP alone is necessary but not sufficient.
- Timing closure costs about 2.5x gates: the ~300 gate claim holds for the naive structure, and the demo version exceeds it. The one-cycle latency holds for both.
- The hardware is not a speedup for the N=1 dense case. The whole 11,172-entry space fits in L1 cache, which the software already exploits at 0.38 ns per access. The hardware value is energy, determinism, and physical embedding, not per-operation speed.
- Toolchain notes: Verilator requires explicit width-matched operands and lint scoping for intentional truncation; icetime needs the iCE40 chipdb discoverable (a Homebrew symlink on macOS); the golden anchor file is a language-neutral contract that can gate a C++ port later.
10 Next Steps
| Track | Step | Dependency |
|---|---|---|
| Phase 3 | Physical Upduino 3.1 bring-up: program the bitstream, wire DIP switch and LED modules per the PCF, capture the demo video | board and modules |
| Phase 4 | OpenROAD standard cell flow on Sky130: timing, area, power report. The flow is set up and runs end to end (synth, floorplan, place, cts, route, finish), measured locally and re-run in the syntagma hw CI job. Two flow bugs were fixed during derisking (virtual clock, kepler-formal LEC crash), see the devlog entry | none |
| chton propagation | Generate the segment store SRAM from chton_sram.py (OpenRAM) and attach it to the decoder |
OpenRAM and PDK |
| Power model | Run the OpenROAD power analysis on the VCD trace from make sim-trace |
Phase 4 |
| Multi-target | Port the demo flow to ECP5 or Artix-7 | board choice |
| Streaming | Pipelined decode at one code point per clock; the acceleration case appears at N greater than 1 | follow-up design |
11 Validation and Status
Verified: the decoder matches the Rust reference on all 11,172 valid syllables in three independent channels, the RTL and the synthesized netlist are formally equivalent over all 2^16 inputs, and the demo closes the 12 MHz board clock with margin (16.79 MHz measured). The compose and distance units are verified the same way over their own input spaces. The segment store behavioral model is verified over all 11,172 slots. The bitstream is generated and ready for board bring-up.
Pending: physical board demonstration and video, the CI confirmation of the OpenROAD standard cell report (measured locally: 4631 um^2, 11.19 ns critical path, 0.897 mW for the demo top; 388 Sky130 cells, 2826 um^2 for the pure decoder), the OpenRAM SRAM macro, and the power model on the VCD activity trace. These are follow-up tracks, not blockers for the current reference.
12 References
The companion whitepapers describe the primitive definition and the software reference implementation; the hardware index carries the measured results summary.