CERN-SSCCS Collaboration Framework

Draft: Collaboration Through Open Science and Open Source Infrastructure

Author
Affiliation

SSCCS Initiative

Abstract

This draft outlines a proposed collaboration between the SSCCS Initiative and CERN. SSCCS is an independent, open-source, public-good science foundation whose work follows the principle that has guided CERN since its founding: science is a public good, and knowledge should remain open. The document presents what SSCCS brings, proposes a concrete first engagement on a ROOT TTree I/O access pattern in production analysis workflows, and outlines how the collaboration would proceed through CERN’s established open-source channels. This draft is the basis for discussion.

Whitepaper
Tech report
Community
Other Formats

Institutional Alignment

SSCCS is an independent foundation operating under the SSCCS Charter, with a primary mission of advancing open structural computing. It is not accountable to any state, party, or corporation, and releases its core work under Apache 2.0. The Charter permits commercial or revenue-generating activities when consistent with its objects, but the foundation’s mission remains public-good oriented.

The proposed collaboration stays within CERN’s stated direction: “unleash the potential of open source for scientific and technological progress beyond high-energy physics.” SSCCS works precisely in that domain: open hardware on RISC-V, structural computing primitives in the Tagma coordinate space, and exhaustive verification tooling.

What SSCCS Brings

The following foundational assets are under active development, with key primitives already validated on silicon simulators and production-grade reference implementations in progress:

Asset Relevance
ExaVerif (exhaustive RISC‑V verification, 33.5M encodings in 29 ms) Exhaustive, deterministic verification for RISC‑V and radiation‑tolerant systems
Chton (zero‑serialization IO materialization) Storage format equals the memory layout; the read path has no serialization step
Tagma (hash‑less coordinate space; decoder ~300 gates) Direct coordinate addressing with a minimal decoder, suited to low‑power and constrained environments
Apache 2.0 release Unencumbered integration with open‑hardware and scientific software ecosystems

A focused starting point

A CMS Open Data analysis whose TTree read path is dominated by per-request I/O overhead.

The case

A production workflow issues hundreds of thousands of singular read requests (totaling about 1.6 GB) over about 14 hours, at an average throughput of tens of KB/s. ROOT’s TTree I/O has been optimized over many years; the measurements in the technical report apply to this specific access pattern and to its cache-disabled baseline. Despite ongoing I/O modernization efforts (e.g., RNTuple), the structural cost remains for this pattern: each read traverses branch hierarchies and deserializes data. Raw disk bandwidth is ample at this scale; the cost is per-request overhead: page cache thrashing from small fragmented reads and cache invalidation under multithreaded access patterns.

The SSCCS Approach

SSCCS addresses this by replacing the read path with coordinate addressing. An event becomes a coordinate resolved to a direct offset, with no branch table, no index scan, and no hash. Filters become spatial operations over coordinate ranges, and Chton materializes the coordinate space onto media with the storage format equal to the memory layout.

The Offer to CERN

ROOT’s I/O interface stays exactly as it is. Researchers keep the TTree API they already know, and their existing analysis scripts run unchanged. Underneath, mapped data is served from the coordinate space instead of the disk. The ecosystem remains intact; only the interior is replaced.

Demonstration

The demonstration has run. The technical report in this directory measures the fork on the CMS Run2016G DoubleMuon NanoAOD first file: the coordinate read path serves every event with one request and zero application-level read system calls where the baseline issues hundreds of thousands of singular reads, and it reproduces the access pattern the case above describes. The measured record now covers the whole event, a block-compressed store, and the event-selected access the production signature is made of, where the baseline pays 683 requests per event against the store’s one. The report states what is measured and what is not: the scan gain, the column-selection crossover, and the conditions that need infrastructure, recorded row by row in the fork’s benchmark harness.

Status

The document has moved from a proposed starting point to a measured one. The fork is open, the report carries the measurement, and the harness, the runbook, and the open fronts are in the fork’s benchmarks/tagma. What remains needs infrastructure rather than method, O(128), multi-terabyte samples, and the remote medium, together with the claim revision that keeps every number scoped to the condition it was measured in.