Data Contracts
AHRI_TRE_RS uses a small set of canonical data contracts so the Rust core, adapters, control plane, examples, and future language bindings do not invent competing formats.
Canonical Formats
| Concern | Contract |
|---|---|
| In-memory tabular data | Apache Arrow RecordBatch values |
| Persisted analytical datasets | Parquet |
| Binary tabular transport | Arrow IPC stream or file |
| Control-plane messages | JSON over HTTP |
The workspace version baseline for Arrow-family crates, Parquet, DuckDB, and
the Rust toolchain is maintained in the root Cargo.toml workspace dependency
table and mirrored in docs/tabular_contract.md.
Arrow At Crate Boundaries
Arrow RecordBatch values are the neutral in-memory tabular boundary. The
ahri_tre_tabular crate provides reusable helpers for schema-validated Arrow
tables, representative typed-column construction, Parquet read/write behavior,
and Arrow IPC stream/file byte helpers.
Python, Julia, and R dataframe types are client-local representations. They are not the canonical wire model and should not drive Rust service or adapter interfaces.
Parquet For Persistence
Parquet is the canonical persisted dataset format for lake-managed tabular data and exported dataset artifacts where a columnar binary format is needed. CSV and newline-delimited JSON may exist as export formats for interoperability, but they do not replace Parquet as the primary persisted dataset contract.
Arrow IPC For Binary Transport
Arrow IPC is the binary tabular transport when JSON is not appropriate. It preserves Arrow schemas and batches across process and language boundaries without treating a language-specific dataframe format as the shared contract.
JSON Control Plane
The control-plane direction is JSON over HTTP. JSON should carry commands, status, diagnostics, metadata envelopes, and machine-readable responses. Large analytical tables should use Arrow IPC or persisted datasets rather than being forced through JSON.
The current protocol, daemon, and CLI crates provide the structure for this
direction. The implemented CLI schema registry now exposes local JSON Schemas
for stable protocol envelopes, shared protocol objects, session, datastore,
daemon readiness, domain, study, governance, asset, datafile, dataset,
lifecycle/delete, ingest, transformation, workflow, dictionary, tag, semantic
model, and bounded CLI-local readiness/lifecycle payloads. Inspect them with
ahri-tre schema list --format json and ahri-tre schema get protocol.dataset.catalog.v2 --format json.
These schemas describe JSON control-plane DTOs owned by ahri_tre_protocol or
bounded local CLI DTOs. They do not replace Arrow IPC or Parquet for analytical
data. Dataset row streams, binary table transfer, and persisted lake/export
artifacts should continue to use Arrow IPC and Parquet where those formats are
the correct contract.
For substantive TRE workflow operations, the JSON payload ownership model is
protocol-first: public request and response bodies belong in
ahri_tre_protocol and are carried in stable protocol envelopes. The local
daemon is an execution and session adapter over those protocol contracts. It
may own local process-control and runtime-only DTOs, but it should not define a
second public JSON grammar for reusable workflow results.
TRE Variable Value Types
The PostgreSQL datastore table public.value_types is the source of truth for
supported TRE variable value types. Runtime code should resolve these values
through the canonical mapping in ahri_tre_types::TreValueType instead of
introducing ad hoc ValueTypeId constants in workflow or adapter code.
| id | datastore string | category | Use |
|---|---|---|---|
| 1 | xsd:integer | scalar | Whole-number values. |
| 2 | xsd:float | scalar | Floating-point numeric values. |
| 3 | xsd:string | scalar | Text values. |
| 4 | xsd:date | scalar | Calendar dates. |
| 5 | xsd:dateTime | scalar | Date and time values. |
| 6 | xsd:time | scalar | Time-of-day values. |
| 7 | enumeration | categorical | One selected vocabulary item per observation. |
| 8 | multiresponse | categorical | Zero or more selected vocabulary items per observation. |
The six XSD-backed types are scalar variable types. enumeration is for
single-response categorical variables represented by a TRE vocabulary, such as
REDCap radio, dropdown, yes/no, and true/false fields. multiresponse is for
multiple-response categorical variables represented by a TRE vocabulary, such
as REDCap checkbox fields.
Boolean and binary are not supported TRE variable value types today. Boolean
source fields may be normalized to an existing scalar or categorical TRE type
by an ingest workflow, and binary concerns belong to tabular transport or file
storage contracts such as Arrow IPC, Parquet, or governed datafiles. They
should not be recorded as public.value_types rows unless a future schema
change explicitly adds them.