Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Data Contracts

AHRI_TRE_RS uses a small set of canonical data contracts so the Rust core, adapters, control plane, examples, and future language bindings do not invent competing formats.

Canonical Formats

ConcernContract
In-memory tabular dataApache Arrow RecordBatch values
Persisted analytical datasetsParquet
Binary tabular transportArrow IPC stream or file
Control-plane messagesJSON over HTTP

The workspace version baseline for Arrow-family crates, Parquet, DuckDB, and the Rust toolchain is maintained in the root Cargo.toml workspace dependency table and mirrored in docs/tabular_contract.md.

Arrow At Crate Boundaries

Arrow RecordBatch values are the neutral in-memory tabular boundary. The ahri_tre_tabular crate provides reusable helpers for schema-validated Arrow tables, representative typed-column construction, Parquet read/write behavior, and Arrow IPC stream/file byte helpers.

Python, Julia, and R dataframe types are client-local representations. They are not the canonical wire model and should not drive Rust service or adapter interfaces.

Parquet For Persistence

Parquet is the canonical persisted dataset format for lake-managed tabular data and exported dataset artifacts where a columnar binary format is needed. CSV and newline-delimited JSON may exist as export formats for interoperability, but they do not replace Parquet as the primary persisted dataset contract.

Arrow IPC For Binary Transport

Arrow IPC is the binary tabular transport when JSON is not appropriate. It preserves Arrow schemas and batches across process and language boundaries without treating a language-specific dataframe format as the shared contract.

JSON Control Plane

The control-plane direction is JSON over HTTP. JSON should carry commands, status, diagnostics, metadata envelopes, and machine-readable responses. Large analytical tables should use Arrow IPC or persisted datasets rather than being forced through JSON.

The current protocol, daemon, and CLI crates provide the structure for this direction. The implemented CLI schema registry now exposes local JSON Schemas for stable protocol envelopes, shared protocol objects, session, datastore, daemon readiness, domain, study, governance, asset, datafile, dataset, lifecycle/delete, ingest, transformation, workflow, dictionary, tag, semantic model, and bounded CLI-local readiness/lifecycle payloads. Inspect them with ahri-tre schema list --format json and ahri-tre schema get protocol.dataset.catalog.v2 --format json.

These schemas describe JSON control-plane DTOs owned by ahri_tre_protocol or bounded local CLI DTOs. They do not replace Arrow IPC or Parquet for analytical data. Dataset row streams, binary table transfer, and persisted lake/export artifacts should continue to use Arrow IPC and Parquet where those formats are the correct contract.

For substantive TRE workflow operations, the JSON payload ownership model is protocol-first: public request and response bodies belong in ahri_tre_protocol and are carried in stable protocol envelopes. The local daemon is an execution and session adapter over those protocol contracts. It may own local process-control and runtime-only DTOs, but it should not define a second public JSON grammar for reusable workflow results.

TRE Variable Value Types

The PostgreSQL datastore table public.value_types is the source of truth for supported TRE variable value types. Runtime code should resolve these values through the canonical mapping in ahri_tre_types::TreValueType instead of introducing ad hoc ValueTypeId constants in workflow or adapter code.

iddatastore stringcategoryUse
1xsd:integerscalarWhole-number values.
2xsd:floatscalarFloating-point numeric values.
3xsd:stringscalarText values.
4xsd:datescalarCalendar dates.
5xsd:dateTimescalarDate and time values.
6xsd:timescalarTime-of-day values.
7enumerationcategoricalOne selected vocabulary item per observation.
8multiresponsecategoricalZero or more selected vocabulary items per observation.

The six XSD-backed types are scalar variable types. enumeration is for single-response categorical variables represented by a TRE vocabulary, such as REDCap radio, dropdown, yes/no, and true/false fields. multiresponse is for multiple-response categorical variables represented by a TRE vocabulary, such as REDCap checkbox fields.

Boolean and binary are not supported TRE variable value types today. Boolean source fields may be normalized to an existing scalar or categorical TRE type by an ingest workflow, and binary concerns belong to tabular transport or file storage contracts such as Arrow IPC, Parquet, or governed datafiles. They should not be recorded as public.value_types rows unless a future schema change explicitly adds them.