References
Standards
Integrations
Other Formats
Spatial coordinate space computing system built on Tagma
Every system today pays a structural cost: finding data. Whether through hash tables, index trees, or pointer-chased lookups, the machine spends cycles on search rather than on computation. synTagma is being developed to eliminate the search layer. In its model, every datum is assigned a fixed coordinate within a deterministic space, and that coordinate serves as its address. Retrieval is designed to be direct arithmetic: given a coordinate, the datum is located by position, with no lookup, no collision resolution, and no index structure to maintain. This shifts data access from a search problem to a geometry problem, removing the indexing overhead that scales with dataset size and access history.
synTagma1 is a spatial coordinate space computing system. Its core primitive, Tagma, is a 16-bit coordinate embedded in a fixed Unicode block (U+AC00–U+D7AF) that replaces hash-based addressing with direct structural addressing. The block sits in the Unicode Basic Multilingual Plane (BMP), so one character is one UTF-16 code unit: one u16, one register, one SRAM word. Every valid 16-bit value is simultaneously a 1-D address (Unicode code point), a 3-D coordinate (Axis 0, Axis 1, Axis 2), and a displayable Unicode character.2 This triple interpretation enables hash-less content addressing with zero collision probability and single-cycle combinational decoding at an estimated ~300 gates3.
synTagma is a full-stack computing system built on this primitive. At the core, Tagma provides the atomic coordinate: a 16-bit value with closed-form three-axis composition and single-cycle combinational decoding. On top of the core, a family of coordinate-space data structures realizes the addressing model in software: a hashless key-value map that stores and retrieves entries by coordinate, a spatial query layer that turns axis algebra into proximity, bounding-box, and distance operations, and a range of direct-address and tree-backed space types that scale from a 22 KB no-allocator array to terabyte-scale mapped spaces. A security layer binds authority, integrity, audit, and channel guarantees to the coordinate structure, and a self-validating compositional encoding (Base11172) provides serialization for storage and transport. The same primitive is realized in hardware as a combinational RTL decoder, verified exhaustively against the Rust reference and carried through FPGA and standard cell flows. Above all of these, the coordination layer extends the arithmetic across physical topologies: each axis of a multi-Coord coordinate can reside on a different node, and because each axis is independently stored and accessed, no distributed consensus is required.
Three structural properties make the coordinate an address:
The Unicode stability policy keeps these constants durable; the general Tagma mechanism extends beyond the BMP with the N-dimensional machinery (CoordPath, base11172, CoordSpace2). No other Unicode character range combines the three structural properties: the block U+AC00–U+D7AF is the only one that does. The full account is in the reference implementations.
The core Tagma coordinate is defined by the composition formula (ISO/IEC 10646) for block U+AC00–U+D7AF:
\[C(i,m,f) = \text{U+AC00} + 588i + 28m + f, \quad 0 \leq i < 19,\; 0 \leq m < 21,\; 0 \leq f < 28\]
Of 65,536 representable 16-bit states, 11,172 satisfy this formula. The remaining 54,364 are structurally invalid and hardware-detectable. Each valid value carries three interpretations: a Unicode code point for flat addressing, a triple-axis coordinate for structural queries, and a Unicode character for display.
| Property | Value |
|---|---|
| Coordinate space | 19 (Axis 0) x 21 (Axis 1) x 28 (Axis 2) = 11,172 valid 16-bit values |
| Invalidity margin | 54,364 of 65,536 states are structurally invalid (hardware-detectable) |
| 1-D interpretation | Unicode code point U+AC00–U+D7AF (character encoding) |
| 3-D interpretation | (Axis 0, Axis 1, Axis 2) coordinate (structural positioning) |
| Display interpretation | Unicode character (debugging aid, zero exposure in the process) |
CoordPath composition extends the address space linearly with Coord count:
| Coords | Axes | Identifier space | Equivalent to |
|---|---|---|---|
| 1 | 3 | \(1.12 \times 10^4\) | Sensor tags |
| 6 | 18 | \(1.94 \times 10^{24}\) | UUID scale |
| 10 | 30 | \(2.69 \times 10^{40}\) | Exceeds 128-bit |
| 19 | 57 | \(1.94 \times 10^{77}\) | SHA-256 scale |
The decoder extracts three axis fields from a 16-bit input in one combinational cycle: range check, field extraction (division by 588 and 28), and axis validation. Total gate count is approximately 300 gates in 28nm – smaller than a single 32-bit multiplier.
Hash-based identity generation costs approximately 10,000 gates and 64-75 cycles per operation. The Tagma decoder replaces this with approximately 300 gates and one cycle – a 30x reduction in hardware cost and a 64-75x reduction in latency.
For larger identifier spaces, Coords compose linearly. Each Coord adds a factor of 11,172 to the addressable space (see benchmark table for addressable space per variant).
CoordSpace is the Rust type family that realizes the Tagma core primitive: a generic direct-address array indexed by coordinate, zero hash and collision. The benchmarks below compare its performance against standard hash maps.4
SHA-256 requires ~10,000 gates and 64-75 cycles per operation, then needs collision resolution and dynamic resizing. UUID generation requires entropy and delivers probabilistic uniqueness. Tagma replaces all of this with a combinational decoder and a 16-bit register. A single Coord covers 11,172 identifiers; six Coords (18 axes) exceed typical distributed system needs; nineteen match SHA-256’s \(2^{256}\) space.
Measured lookup latency: 0.39 ns for a single-Coord native CoordSpace (dense array, no allocator) vs 227 ns for SHA-256 (582x). Native CoordSpace is flat across all N: 0.39 ns at every depth, because every Coord resolves to a single array load regardless of Coord count. The tree fallback (CoordSpaceN) scales linearly with N: 2.69 ns at N=3, 58.6 ns at N=19, because each level requires a heap dereference and enum match. Recursive depth is bounded by schema, not by data volume: \(10^4\) and \(10^{77}\) entries both cost \(N\) dereferences in the fallback path, while the native dense path costs a constant 0.39 ns.
Nonexistent prefix lookup: CoordSpace 1.65 ns (structural, navigates to the branch and returns None) vs HashMap 23.05 ms (14.0Mx, full scan of 10M entries; HashMap has no structural prefix index). Sparse get at 10M entries: CoordSpaceN2 completes all 10M operations in 44.9 ms vs HashMap 1.05 s (23.4x).
This is the elimination of hashing itself, not replacing HashMap as a storage, which is one of the fastest general-purpose hash-based storages by C-grade machine code. Tagma matching or exceeding this baseline is incidental: HashMap degrades linearly while Tagma does not.
SHA-256 lookup costs 227 ns; the tree fallback (CoordSpaceN) reaches \(2^{256}\) at 19 Coords for 58.6 ns (3.9x). The native dense path (CoordSpace / CoordSpace2 / CoordSpaceM3) holds at a flat 0.39 ns.
Address space grows with N; the tree fallback lookup cost scales as O(N), while the native dense path is O(1) flat. Tagma recursion k=1 reaches \(10^{231}\) identifiers (SHA-512 space × \(10^{77}\)) at 171 ns, exceeding every hash system at a fraction of the cost.
| Metric | SHA-256 | CoordSpace (N=1) | CoordSpace2 (N=2) | CoordSpaceM3 (N=3) | CoordSpaceN (any N) |
|---|---|---|---|---|---|
| Latency per lookup (ARMv8.4-A Firestorm) | 227 ns | 0.39 ns | 0.39 ns | 0.40 ns | 0.94-40.8 ns |
| Allocation | – | Inline 22 KB | Heap 119 MB | Mmap 1.27 TB | Sparse tree |
| Identity size | 32 bytes | 2 bytes | 4 bytes | 6 bytes | \(2N\) bytes |
| Addressable space | \(2^{256}\) | \(1.12 \times 10^4\) | \(1.25 \times 10^8\) | \(1.39 \times 10^{12}\) | variable |
| Collision | probabilistic (\(2^{-128}\)) | deterministic zero | deterministic zero | deterministic zero | deterministic zero |
| Tagma | – | Complete | Complete | Complete | Fallback |
Spatial query: CoordSet bitwise AND resolves compound axis filters at 329 Melem/s, 137x faster than HashMap scan. Edge: CS2 sparse get sustains 23.4x at 10M entries; CoordSpaceN19 get shows the 19-dereference cost (0.50x); drain is 0.72x on the full space.
Bulk operations: CoordSpace outperforms HashMap by 14.6-17.3x across all operations on the full 11,172-entry space. The single-get microbenchmark isolates the per-operation cost: 0.82 ns vs 8.50 ns.
Stress test: under 500,000 interleaved insert, get, remove, and update operations, CoordSpace completes in 3.64 ms vs HashMap 12.2 ms. Deep tree: CoordSpace2 and CoordSpaceM3 (dense, N=2 and VM N=3) reach 0.05 µs for 100 gets (0.39 ns per access), CoordSpaceN19 incurs tree traversal cost (7.04 µs), and HashMap stays at 3.53 µs. Nonexistent key lookup favors dense encoding (0.39 ns) over tree depth (2.30 ns) and hash miss (20.1 ns).
Tagma assigns every point in a geometric space a structural address that is simultaneously a coordinate, an identifier, and a computation target. HashMap stores values by hashing keys by comparison. Querying this space is spatial computation: axis projection, set membership, proximity, and coordinate slicing are arithmetic operations, not index scans. The figures above measure the consequence: HashMap degrades with data volume; the coordinate space does not.
Cross-validation from hardware verification: ev (ExaVerif) confirms the same structural advantage on real RISC-V instruction encoding spaces. At CVA6 CV-X-IF scale, structural enumeration verifies the full 33M-combination encoding space in 18.5 ms versus 18.0 s for the standard pipeline, a 973x speedup; the Ibex RV32IMCB Report space shows the same pattern at 524,288 combinations, 79x.
Three allocation strategies produce distinct resource profiles. CoordSpace2 (dense) preallocates the full 11,172 x 11,172 grid as a flat array: 119 MB, one allocation call, one cache miss per lookup. Cost is fixed regardless of occupancy: at 10K entries the 11900 B/entry overhead is high, but at 10M entries it drops to 11.9 B/entry. CoordSpaceN2 (tree) allocates one 44 KB leaf node per written prefix (11,172-slot). Memory scales with prefix count, not entry count: 10001 allocations for 10K entries across a handful of prefixes, 45 B/entry at this density. HashMap allocates per-entry: 10 million insertions produce 10 million allocation calls, each with malloc overhead, bucket resizing, and rehashing.
The practical consequence: dense eliminates allocation entirely but pays a fixed memory floor. Tree memory is fixed at prefix-creation time and independent of per-prefix density. HashMap’s footprint and allocation cost grow with every entry and exceed both coordinate-space strategies at scale.
Secondary benchmarks confirm the same scaling pattern across all depths. CoordSpaceN2 insert/get at 1,000 entries completes in 791 µs (insert) and 4.99 µs (get). At 100,000 entries with 1,000 unique prefixes, CoordSpaceN2 insert allocates nodes in 15.6 ms while HashMap inserts in 6.6 ms without preallocation; get is 48.2 µs vs 32.1 µs. These non-core paths follow the expected tradeoff: tree node allocation overhead at small scale inverts at large scale where HashMap’s per-entry cost compounds.
The Tagma coordinate space is exhaustively enumerable: 65,536 possible 16-bit values, of which exactly 11,172 satisfy the composition formula. A verification harness generates all 65,536 inputs, applies the decoder specification, and records every result. The bijection is verified by enumeration: no collisions, no unassigned values within the block. Every implementation, hardware or software, can be verified against the same ground truth in milliseconds on any commodity system.
A SHA-256 engine cannot be exhaustively verified across its full input space. The Tagma coordinate space can, because it is bounded and its formula is closed-form.
| Method | Identifier size | Generation cost | Collision | Lookup |
|---|---|---|---|---|
| Pointer | 32-64 bits | zero | none | direct |
| Hash (SHA-256) | 256 bits | ~10K gates, 64-75 cycles | probabilistic | hash table + resolution |
| UUID | 128 bits | entropy-dependent | probabilistic | hash table |
| CAM | per-bit comparison | 9-16 transistors/bit | none | associative |
| Tagma | 16 bits | 1 cycle (combinational) | none (formulaic) | direct (coordinate = address) |
| Domain | Hash Role | Tagma Replaces? |
|---|---|---|
| ID generation | hash(data) for unique identifier | Yes |
| Hash map key | hash(key) for bucket index | Yes |
| Content addressing | data hash as storage address | Partial |
| Integrity verification | data hash for tamper detection | No (retain SHA-256) |
| Digital signatures | hash-then-sign for non-repudiation | No (retain Ed25519) |
| Key derivation | HKDF for key expansion | No (retain HKDF) |
The entries marked No are not limitations; they reflect a deliberate composition strategy. An identifier can be generated by SHA-256 and then encoded as Tagma coordinates at write time with near-zero overhead – a single modulo operation per hash block. From that point forward, every read operation uses Tagma’s native O(1) direct access regardless of how the identifier was originally produced. The cryptographic cost is paid exactly once at write time; all subsequent accesses benefit from the structural address space.
This composition pattern extends beyond cryptography to every domain in the table. Each domain can adopt Tagma at its own pace – replacing the hash step where possible, wrapping it where necessary – while always reading at Tagma speed. The result is a universal addressing layer that different applications (cryptographic content addressing, distributed node identification, sensor networks, real-time object tracking) share without compromising their domain-specific requirements.
The coordination layer extends Tagma’s arithmetic to physical topologies without modifying the core:
Recursive coordinate space. The composition formula admits unbounded \(k\) levels of recursion. A 19-Coord sequence at \(k=0\) occupies the SHA-256 scale; at \(k=1\) each of its three axes is itself a full 19-Coord sequence.
Self-routing. A Coord sequence carries its own address. Given a sequence and current recursion depth, any node determines whether each sub-axis is local or remote by applying the topology function \(\phi_k\). No routing table lookup is required.
Axis fungibility. The three axes carry no intrinsic semantics. Axis 0 may represent “region”, “shard”, or “timestamp” depending on deployment. The coordinate arithmetic – composition, decomposition, linearisation – is invariant.
The Tagma coordinate space is implemented as a Rust library providing:
The CoordSpace family follows a three-tier design that scales from embedded to datacenter:
| Variant | Allocation | Access | Latency | Tagma |
|---|---|---|---|---|
| CoordSpace (N=1) | Inline 22 KB | Single load | 0.39 ns | Complete |
| CoordSpace2 (N=2) | Heap 119 MB | Single load | 0.39 ns | Complete |
| CoordSpaceM3 (N=3) | Mmap 1.27 TB | Single load | 0.40 ns | Complete |
| CoordSpaceN<N> (any N) | Sparse tree | N dereferences | 0.94-40.8 ns | Fallback |
&[Coord] runtime paths. Used as fallback when the dense variants exceed system memory.Embedded systems. The 11,172-identifier space fits in a 22 KB no-allocator array: one load, no hashing, no collisions.
LLM inference cache. KV caches indexed by token prefixes use CoordPath-based direct access: production cache sizes (\(10^4\)–\(10^7\)) are covered by 2–4 Coords, with zero hash computation.
Graph and multi-dimensional query. Each node maps to a Coord, each edge type to a CoordSet. Adjacency reduces to a bitwise AND over 175 machine words – no index intersection.
General-purpose addressing. Replaces UUIDs, hash keys, and sequence numbers with shorter, faster, deterministic identifiers.
Tagma replaces hash-based identity generation and addressing. SHA-256 remains for signatures, Merkle proofs, and integrity verification. Encryption, authentication, and key derivation are outside the primitive’s scope. The two strategies compose: SHA-256 output encoded as 19 Coords is more readable than 64 hex characters while preserving \(2^{-256}\) collision probability.
synTagma is built on open international standards and runs on open ISA (RISC-V), unencumbered by proprietary licensing or regional dependencies. The Tagma primitive is a neutral, sovereign coordinate space for the global computing landscape.
This block was allocated in Unicode 2.0 (1996).5 Its encoding formula provides three independent axes, Axis 0, Axis 1, and Axis 2, the coordinate ranges used by Tagma. This structure is analogous to an apartment building numbering system: given “101”, you know the floor and room without a central directory. Tagma leverages this pre-existing coordinate space as a universal address space, eliminating the need for hash-based directories.
The primitive family and its application domains are organized as follows:
| Document | Description |
|---|---|
| Tagma | Primitive whitepaper: coordinate space, decoder, coprocessor, compliance, benchmarks, SEU analysis |
| TagmaMap | Hashless Key-Value Storage (CoordSpace type paper) |
| TagmaGeo | Spatial Query Layer on Structural Coordinate Spaces (CoordCube type paper) |
| TagmaMatrix | Coordinate-addressed rank-2 matrices, the integer product, and the wire form (Matrix type paper) |
| Serialization | Base11172: human-readable, self-validating serialization |
| Identification | Content-addressable identity without hash functions (application domain) |
| Security | Integrity, authorization, audit, and non-repudiation on a structural coordinate space (application domain) |
| Signal | Structural coordinate transmission: the coordinate is the message (application domain) |
| Hardware | Hardware implementation of the Tagma primitive: RTL, synthesis, FPGA demo, chton SRAM |
| Memory | Three-dimensional SRAM decode and the chton SRAM segment store |
| Vec (forthcoming) | Vector processing over structural coordinates |
| Tagma benchmarks | 51 microbenchmarks across 12 criterion groups: identity generation, spatial query, edge cases at 10M entries, deep trees, mixed-operation stress tests |
| Ibex RV32IMCB Report | Exhaustive RISC-V verification using structural enumeration: 79x speedup |
| CVA6 CV-X-IF Verification | At 33M scale, structural enumeration achieves 973x speedup |
For inquiries or research collaboration: syntagma@ssccs.org.
© 2026 SSCCS Initiative — Open-source computing systems initiative building a computing model, software compiler infrastructure, and open hardware architecture.
synTagma is an independent open-source project (Pre-release, Apache 2.0) under the SSCCS Initiative.↩︎
The ranges 19, 21, and 28 derive from the Unicode block U+AC00–U+D7AF, which encodes the compositional writing system. The composition formula is defined in ISO/IEC 10646.↩︎
Pre-silicon estimate based on the gate-level specification; actual gate count and cycle timing will be confirmed through standard-cell synthesis and static timing analysis with the coprocessor (e.g. RISC-V XIF).↩︎
All benchmarks were run with cargo bench, criterion on ARMv8.4-A Firestorm. Code: Github↩︎
The constants 19, 21, and 28 correspond to the number of initial consonants, medial vowels, and final consonants in Hangul (invented 1443), the compositional writing system encoded in block U+AC00–U+D7AF.↩︎