Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

CLI manual

Governed content control documents are available through ahri-tre schema get protocol.content-transfer.v1. This schema describes admission requests, immutable input identities, exact effective budgets, deployment limits, and terminal evidence. Content is delivered incrementally; admission success alone does not establish completion. See the content transfer contract for destination and C payload lifetime semantics.

For an explicit Dataset destination:

ahri-tre --session analysis --study demo dataset data \
  --dataset visits --format arrow --to visits.arrow

Direct queries declare each input’s owning Study or immutable reference; they do not accept the global --study selector:

ahri-tre --session analysis query \
  --inputs '[{"alias":"visits","asset":{"kind":"name","study":{"kind":"name","name":"demo"},"name":"visits","asset_type":"dataset"}}]' \
  --sql 'SELECT * FROM visits LIMIT 100' --format parquet --to sample.parquet

The shared request budgets are --max-payload-bytes, --max-decoded-file-bytes, --max-rows and --max-transfer-seconds. Omitted values use deployment defaults. Explicit positive values retain their exact value within the advertised ceiling; invalid values fail before work. --limit or SQL LIMIT selects rows independently of these service budgets. Dataset commands accept --version; a complete Asset-version reference also pins content. Datafile export supports --no-decrypt, --no-decompress and logical-byte recompression with --compress.

Use --to - for payload-only stdout. For file destinations, existing files are preserved unless --overwrite is explicit, and even then replacement follows verified completion. An error on stdout can leave partial bytes and returns a nonzero exit status. Disclosure identity and notices go to stderr.

All commands accept --config PATH for explicit document selection and most reporting commands accept --format text|json. Explicit selection is strict.

Top-level command groups are:

  • version, doctor, schema, and completion;
  • config for schema, initialization, validation, Effective rendering, preflight, and Client rendering;
  • secrets for privileged Managed-secret administration;
  • daemon for the Managed runtime lifecycle;
  • session and datastore for authenticated protocol operations;
  • governed metadata and data groups such as domain, study, asset, datafile, dataset, variable, vocabulary, tag, entity, entity-relation, transformation, and ingest.

Use ahri-tre <group> --help for the exact installed grammar. Removed direct database profiles, dotenv files, endpoint overrides, cached-token opens, passfiles, Lake moves, and adoption/reset commands are intentionally absent.

Detached Dataset materialization

ingest dataset from-datafile waits for completion by default. To receive an operation ID as soon as the server accepts the work:

ahri-tre --session analysis --study StudyA ingest dataset from-datafile \
  --domain HDSS --dataset observations --source-asset source_csv \
  --description "Materialize observations" --no-wait --output-format json
ahri-tre --session analysis operation list --scope datastore --status completed --limit 25 --format json
ahri-tre --session analysis operation get OPERATION_ID --event-limit 25 --format json
ahri-tre --session analysis operation cancel OPERATION_ID --format json
ahri-tre --session analysis operation result get OPERATION_ID --format json

Acceptance JSON contains kind: "operation"; the normal successful wait returns kind: "completed" with the materialization result. Text detach output names the accepted operation. Interrupting the CLI wait leaves accepted work running while its original Session stays available. Failed or cancelled waits exit nonzero; successful inspection of a failed operation exits zero. Current authorization is required for each inspection, including from an explicitly reopened Session.

--source-asset accepts the returned canonical Asset reference or immutable version reference as JSON, as well as a Datafile name. Source and Output must belong to the same resolved Study. A separate --source-version must agree with a pinned reference. Study/Domain names and references can be mixed; equivalent selectors retain the same operation for the same idempotency key.

Existing governed Datafiles can be parsed as CSV, JSON, Arrow IPC, Parquet or XLSX, with the existing parser options. An explicit --format must agree with the Datafile’s declared format. This parser support does not extend the current CSV-only client upload transport; broader acquisition is tracked separately. The derived Dataset remains High and uses the existing version allocation.

This producer requires configured bounded execution and independent lifecycle metadata access. Configured runtimes advertise the four implemented operation commands.

Cancellation returns the current operation summary. cancel_requested means the request was accepted; the operation may still be waiting for an adapter call or cleanup. Inspect it until it becomes cancelled, completed, or failed. Repeated pending cancellation returns the same state. Terminal operations and final admission return conflict. Cancellation requires current owner, Session, Datastore, and resource authority; an invisible ID behaves like an absent ID.

Operation history requires explicit --scope datastore or --scope session. Session scope currently returns an empty collection. Datastore scope lists only the current owner’s authorized work, newest first, with descending operation ID breaking creation-time ties. Repeat --status to include several states; --operation-kind ingest.dataset.from_datafile, --created-after, and --created-before narrow history. Time bounds are exclusive RFC 3339 timestamps.

List --limit and get --event-limit accept 1–500 (default 100). Continue list pages with --cursor and event pages with --event-cursor, copying the returned next_cursor and retaining the query context and filters. Events sort by ascending sequence. New operations do not enter an existing list traversal; statuses and authorization are checked on every page. Cursors are opaque and expire when the Trusted runtime restarts; start a fresh traversal then.

Progress reports only optional stages. A terminal operation does not establish Dataset availability for another write: retained output reservations may still require cleanup. History retention is enforced on every read.

Operation retention

Terminal summaries, events and typed receipts expire exactly 30 days after finished_at; active work has no deadline. Get, event and result reads return not-found at operation expiry, and history omits expired rows before pagination and cursor lookahead, even when physical cleanup is delayed. Completed status is preserved if its result expires early, is removed, or becomes unavailable. The summary’s result reference and operation_result_unavailable detail distinguish expired, removed and unavailable; current source and output authority still apply. Successful result receipts include retention and availability alongside kind and data. Status and result text output display these protocol fields.

Idempotency protection lasts while active and until both acceptance plus 24 hours and the terminal operation deadline have passed. If public history expires inside that minimum key window, a retry returns not-found without executing new work. Expired keys may be reused only after the normal output-reservation checks. Startup and explicit Session opening perform bounded metadata housekeeping, at most 100 expired operations per pass. Housekeeping never releases reservations or deletes private cleanup evidence, Datafiles, Dataset versions, Transformation provenance or audit records. No public prune command or retention setting exists.