Governance And Provenance
Governance and provenance are first-class parts of the AHRI_TRE_RS model. They are not presentation-layer concerns and should not be bypassed by the CLI, daemon, examples, or language bindings.
Governance Separation
The governance model separates four concerns:
- authentication: who the caller is
- datastore entry: whether the authenticated caller may open the datastore
- study authorization: whether the caller may read, write, or administer a particular study
- DUO restrictions: what data-use restrictions apply to a study or asset version
DUO metadata informs policy and review, but it does not grant access by itself. Study access grants, custodianship, and explicit policy checks remain separate concepts.
The Authenticated TRE user is the user identity proven by a live Session’s authentication context. Auditable destructive operations record this identity as their actor when it is available. It is not the same thing as the PostgreSQL current user, a DuckLake catalog user, or study custodianship. Authenticated users with study access may perform eligible study-scoped content lifecycle deletes; deleting a study itself remains a study administration action governed by custodianship or administrator capability. Unused semantic catalog deletes require an authenticated actor and dependency-safe targets, and study-scoped references block deletion rather than becoming authorized through custodianship.
Study Custodianship
Study custodianship is the workflow-facing accountability model for study authorization. A study has one primary custodian and may have delegate custodians. Custodians can list visible study access grants and custodian assignments, and can grant or revoke ordinary study access through the application/database workflow path. They should not write raw governance tables directly.
The primary custodian carries ownership accountability for the study. Only the current primary custodian can add delegate custodians, remove delegate custodians, or transfer primary custodianship. Delegates can help administer ordinary access grants, but they cannot change primary ownership or custodian membership.
Adding a delegate custodian automatically grants that delegate study access when
the access row is missing. Removing delegate custodianship removes only the
study_custodians row; it does not revoke study_access. This keeps
administrative responsibility separate from research access. If a former
delegate should also lose study access, revoke the access grant explicitly.
Primary transfer is a handoff between custodians. The transfer target must already be an existing delegate custodian, so ownership cannot be transferred directly to an arbitrary user. After transfer, the new custodian is primary, the previous primary remains a delegate, and study access is preserved for both principals.
Study access and delegate targets must be existing datastore principals. The PostgreSQL workflow functions validate that the target is an ORCID-backed login role and that it inherits a datastore-entry group role. This prevents governance metadata from naming users who cannot enter the datastore.
study_access uses an open-access fallback: zero rows for a study means the
study is public/open to connected TRE users. One or more rows means the study is
restricted to the listed principals. The grant workflow protects the
public-to-restricted transition by inserting the current primary custodian and
the requested target in the same database operation when the first non-primary
access grant is added to an open study. The revoke workflow also prevents the
primary custodian from being removed from a still-restricted study, while still
allowing the final access row to be removed deliberately to return a study to
public/open access.
Governed Querying
Lake queries are not just SQL execution. The app-layer governed query path requires an authorizer before DuckDB statements are prepared. This keeps access decisions visible at the workflow boundary and prevents the lake adapter from becoming a permissive shortcut around metadata policy.
Provenance Model
AHRI_TRE_RS records lineage around data movement and transformation. The current model uses transformation records plus transformation input/output links for workflow provenance.
Provenance appears around operations such as:
- file ingest
- file-to-dataset conversion
- SQL-to-dataset ingest
- dataset and datafile export
- transformation workflows
- archive-first deletion flows as they are implemented
The transform_assets concept should remain an explicit higher-order workflow
concept. It should not be collapsed into generic CRUD operations, because the
workflow boundary is where governance, inputs, outputs, and review semantics
are easiest to preserve.
Transformation Source Provenance
Transformation records can carry source-code provenance in addition to input/output lineage:
file_path: the workflow source file, script, or notebook responsible for the transformationrepository_url: the Git repository containing that source when discoverablecommit_hash: the Git commit used when discoverable
The shared application provenance path enriches missing Git fields before
transformation records are persisted. If a transformation has a file_path, the
app layer uses it as the source context for Git discovery, reads HEAD, reads
the origin remote, normalizes SSH-style remotes to HTTPS-style URLs, and trims
a trailing .git. Explicit caller-provided values are preserved and are not
overwritten.
This enrichment is intentionally best effort. Non-Git execution environments,
packaged examples, missing remotes, or unavailable git commands do not block
the workflow; the transformation is recorded with the provenance fields already
available.
There is an important Rust-specific nuance. Unlike the Julia
git_commit_info helper, Rust async service functions do not reliably expose the
top-level caller source location through #[track_caller]. Relying on an async
app-service frame would risk recording an internal ahri_tre_app source file
instead of the workflow that caused the transformation. The mitigation is to
record the top-level workflow source path in file_path when constructing the
transformation. The shared app layer then derives repository_url and
commit_hash from that path while preserving the workflow-level file_path.
Documentation Status
Some governance and provenance foundations are implemented, but not every public workflow is complete. When a page describes planned CLI, daemon, binding, or example behavior, it should say so directly and link back to the ordered backlog or issue tracker.