Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

AHRI_TRE As A Trusted Research Environment

AHRI_TRE_RS is being built as the technical core of a trusted research environment (TRE), also called a secure data environment (SDE). Its purpose is to let approved research happen close to sensitive data while preserving the governance, provenance, and operational controls needed to keep that work accountable.

In this documentation, “trusted” does not mean that the software asks operators or researchers to trust it blindly. It means the system is designed so trust can be made concrete: users authenticate through a known identity path, study access is checked at workflow boundaries, data movement is recorded as provenance, datasets use stable tabular contracts, and public outputs are separated from credential-bearing or sensitive operational state.

AHRI_TRE_RS is therefore not a whole institutional governance programme by itself. Ethics review, data-owner approval, researcher accreditation, disclosure review, and incident response still belong to the institution operating the environment. The system supplies the technical spine that lets those institutional controls be expressed consistently in software.

The Five Safes

The Five Safes framework is a common way to reason about controlled research access to sensitive data. UK Data Service summarises the five dimensions as safe data, safe projects, safe people, safe settings, and safe outputs, and describes the framework as a set of principles for safe research access. The Office for National Statistics also frames the model as a way to maximise research use while keeping data secure and preventing identification.

The five dimensions are deliberately interdependent:

SafeQuestionAHRI_TRE support
Safe projectsIs this a legitimate, approved use of the data?Study, domain, DUO restriction, tag, and provenance metadata give approved work a durable place in the datastore model.
Safe peopleIs the user known, trained, and authorised for this work?OAuth/OIDC authentication, datastore-entry checks, study access grants, and custodianship workflows separate identity, entry, access, and administration.
Safe settingsIs the work happening in a controlled environment?PostgreSQL metadata access, DuckDB/DuckLake analytical execution, Trusted-runtime Sessions, selected Execution profiles, and adapter boundaries keep live capabilities inside controlled runtime paths.
Safe dataIs the data prepared and governed for the intended level of access?Managed ingest, typed variables, vocabularies, DUO restrictions, Arrow RecordBatch boundaries, Parquet persistence, and governed query paths keep data treatment visible.
Safe outputsAre released results checked so they do not disclose sensitive information?Export workflows, transformation provenance, redacted diagnostics, secret-safe protocol output, and leak-canary testing create enforceable public-output boundaries. Human disclosure review remains an operational step.

This framing is useful because it prevents over-reliance on any single control. Anonymisation or de-identification is not enough on its own. A TRE needs an approved purpose, trusted users, controlled execution, appropriate data treatment, and governed release of outputs.

AHRI_TRE And Data Visiting

The article “Data visiting governance: a conceptual framework” uses “data visiting” for a model where analysis occurs in the data provider’s controlled computing environment rather than moving the data to the researcher. It argues that data visiting governance should be configurable across several dimensions, including researcher autonomy, data location, data visibility, shared-data nature, output governance, trust and control model, and auditability and traceability.

AHRI_TRE_RS fits naturally into that discussion. It is not only a file store or a metadata catalogue. It is a workflow-oriented environment in which technical choices act as governance levers:

Data-visiting dimensionAHRI_TRE design response
Researcher autonomyThe CLI, daemon, future HTTP control plane, and future language bindings can expose different workflow surfaces while routing through the same app-layer policy checks.
Data locationAnalytical data stays in a Lake mount seen by the runtime as the configured container-visible Lake location, with DuckDB/DuckLake access behind the lake adapter. Ordinary identity-bound opens resolve persisted lake facts from datastore binding metadata.
Data visibilityGoverned query workflows can expose datasets, previews, exports, or transformation outputs without requiring raw database or lake credentials to become user-facing interfaces.
Nature of shared dataStudies, domains, variables, vocabularies, entity links, datafiles, datasets, and DUO restrictions give the datastore enough structure to distinguish source files, governed datasets, semantic metadata, and export artifacts.
Output governanceDataset and datafile export paths are workflow operations that can record provenance and enforce redaction rules at public boundaries. Disclosure review policy can be placed around those exported artifacts.
Trust and control modelThe system is layered so institutional control remains in metadata and workflow policy, while infrastructure details stay behind adapters. This supports central TRE operation today and leaves room for future federated or remote-query surfaces.
Auditability and traceabilityTransformation records, input/output links, Git source provenance, session metadata, redacted diagnostics, and schema/version evidence make work inspectable after the fact.

The key design implication is that AHRI_TRE should be read as a configurable research environment, not a single access mode. A high-trust internal analyst may need interactive session-backed commands. A batch workflow may need stateless automation with explicit inputs and JSON output. A future remote researcher may need a narrow protocol or HTTP surface. These modes should vary the amount of autonomy and visibility without changing the underlying governance and provenance model.

How The Technical Design Supports Trust

Separate Governance Concerns

The governance model separates authentication, datastore entry, study authorisation, and DUO restrictions. This separation matters for both the Five Safes and data visiting. A person can be authenticated but not authorised for a study. A study can carry data-use restrictions without those restrictions being treated as access grants. A custodian can administer access without becoming the only representation of research legitimacy.

The current model records study access grants and study custodianship in PostgreSQL metadata. Study custodians can manage ordinary access grants through workflow paths, and destructive or lifecycle operations can record the authenticated TRE user as the actor. This creates an audit trail around who did what, under which study context, instead of leaving those facts in terminal history or informal operating notes.

Keep Data Near The Provider

AHRI_TRE_RS uses a dual-store runtime:

  • PostgreSQL stores metadata, governance, access control, provenance, and domain structure.
  • DuckDB plus DuckLake stores and queries analytical datasets in the mounted lake filesystem.

This structure supports data visiting because analytical work can happen inside the provider-controlled runtime. The researcher or workflow interacts with controlled commands, protocol requests, Arrow batches, or approved exports rather than receiving broad direct access to every backing store.

The system deliberately treats cross-store work as workflow orchestration rather than pretending PostgreSQL and DuckLake are one distributed transaction manager. That keeps failure handling, cleanup, and provenance visible.

Use Stable Data Contracts

Trusted research environments need predictable boundaries between tools. AHRI_TRE_RS uses:

  • Arrow RecordBatch values as the canonical in-memory tabular model.
  • Parquet as the canonical persisted analytical dataset format.
  • Arrow IPC for binary tabular transport.
  • JSON over HTTP for control-plane messages.

These contracts help safe data and safe settings at the same time. They allow Python, R, Julia, CLI, daemon, and future HTTP surfaces to share the same core tabular and control-plane shapes without making any one client dataframe model the system of record.

Keep Infrastructure Behind Adapters

PostgreSQL metadata access is behind the libpq-based metadata adapter. OAuth and OIDC logic are separated from libpq wiring. DuckDB, DuckLake, and Lake location logic are behind the lake adapter. Runtime configuration is typed and explicit.

Those boundaries are safety controls as much as engineering controls. They make it harder for a CLI command, language binding, or future HTTP handler to bypass policy by reaching directly into PostgreSQL or DuckLake. They also help operators reason about where credentials, live handles, Restricted local references, and query execution can appear.

Treat Public Output As A Boundary

Safe outputs are not just final tables in a publication. In a software system, public output includes CLI text, JSON responses, logs, diagnostics, generated artifacts, archive manifests, provenance summaries, exported metadata, and error messages.

AHRI_TRE_RS has a shared sensitive-material policy for this boundary. Secret material is write-only at credential ingress points and must not appear in public responses. Safe credential metadata can be shown when useful. Restricted local references are redacted or bounded by context. Leak-canary tests give this contract executable evidence.

This does not replace human disclosure control for research results. It does make the software boundary safer by default, so operational output review is not fighting accidental credential or path leakage at the same time.

What This Means For Readers

When you read the rest of this book, the details should connect back to this TRE/SDE model:

The short version is: AHRI_TRE_RS supports the Five Safes by making governance, identity, controlled execution, data treatment, provenance, and safe output boundaries part of the system architecture rather than optional conventions around it.

References