Workflow Documentation Pattern
Use this pattern for every example workflow page. The goal is consistency: a reader should be able to compare HDSS, REDCap, future import workflows, and future analysis workflows without guessing which parts are implemented.
Required Header
Start with the workflow name and a short status block.
# <Workflow Name>
Status: <complete | partial | provisioning-only | smoke-test-only | roadmap>
Implementation owner: <crate, example, command, or planned surface>
Backlog context: <ordered backlog item or local issue path>
Choose the narrowest status that is true today:
| Status | Meaning |
|---|---|
complete | The documented user-facing workflow exists and has ordinary validation. |
partial | Some workflow steps exist, but the page must name missing command groups, adapters, or product behavior. |
provisioning-only | The code prepares infrastructure or sample state but does not yet expose the whole workflow to users. |
smoke-test-only | The behavior is proven through an opt-in test or fixture path, but is not yet a normal documented user command. |
roadmap | The page describes intended behavior and must not read like current product capability. |
Purpose
Explain what research or operational job the workflow performs. Keep this section user-facing and brief, but name the AHRI_TRE concepts involved: domain, study, datastore session, metadata store, lake, asset, datafile, dataset, variable, vocabulary, transformation, provenance, governance, or export.
Current Implementation Status
List what is implemented, what is partial, and what remains planned. Link to source material rather than duplicating backlog text.
Recommended links:
docs/ahri_tre_rust_ordered_backlog.mdfor product and architecture status- local issue files under
docs/issues/for recent implementation slices - example crate README files under
examples/ - crate README files under
crates/ - focused smoke tests when the behavior is intentionally opt-in
Use direct language such as “implemented in the app layer”, “available only as an example binary”, “validated by an opt-in live smoke test”, or “planned for a future CLI/daemon surface”.
Prerequisites
Document the required runtime context before commands appear.
Typical prerequisites:
- a bootstrapped PostgreSQL metadata datastore
- a Datastore binding with its canonical Lake location available to the runtime
- DuckDB/DuckLake support in the development container or deployment target
- local fixture files, external exports, or source-system credentials
- an existing domain or study ID when the workflow does not create one
- live-service opt-in flags for smoke tests
If a workflow needs datastore schema compatibility, say which
datastore schema-status state is required. Prefer datastore schema-plan and
datastore schema-migrate for supported metadata-only upgrades; require
recreation only when status is unsupported or the workflow needs a future
storage-aware migration.
Authentication Modes
State which datastore authentication modes the workflow supports.
Use this checklist when relevant:
- direct PostgreSQL credentials
- interactive OAuth
- stored OAuth artifact
- injected bearer token
- daemon-backed session
- unauthenticated local fixture parsing
- external source-system credentials, such as REDCap API tokens
Do not imply that all modes are equivalent unless the implementation proves it. If a workflow has selected-session behavior, say whether all metadata and lake operations use the selected datastore session.
Inputs
List workflow inputs in a table.
| Input | Required | Source | Notes |
|---|---|---|---|
--data-dir | Yes | Local filesystem | Example: fixture directory containing workflow CSV files. |
| Lake location | Yes | Persisted Datastore binding | Canonical location selected by the binding; never substitute a Restricted local reference or environment authority. |
<domain_id> | Sometimes | Metadata store | Required when the workflow attaches outputs to an existing domain. |
Use exact flag names, logical identifiers, document fields, files, table names, or ID types when they exist. Treat repository-only runner variables as fixture inputs, never product configuration. Use placeholders only for roadmap pages.
Commands
Show commands only for implemented entry points. Prefer copy-pasteable commands that match the repository’s current package names.
cargo run -p <package> -- <flags>
For smoke tests, keep the opt-in variable visible:
RUN_LIVE_<WORKFLOW>_SMOKE=true cargo test -p <package> <test_name> -- --nocapture
If a workflow will eventually be run through the CLI or daemon but that surface is incomplete, place those commands under a “Planned Commands” subsection and label them as planned.
Outputs
Document expected outputs by store and by concept.
| Output | Store | Expected result |
|---|---|---|
| Metadata records | PostgreSQL | Domains, studies, assets, variables, vocabularies, or provenance rows. |
| Managed files | Lake location | Preserved source artifacts, staged files, or exported datafiles. |
| Datasets | DuckLake | Materialized analytical datasets with registered metadata. |
| Diagnostics | CLI, test output, or logs | Health checks, skipped live tests, validation warnings, or provenance summaries. |
Separate ordinary outputs from cleanup artifacts, debug logs, and temporary staging paths.
Validation
Describe how maintainers prove the workflow still works.
Recommended validation sections:
- fast unit or fixture tests that run without live services
- integration tests that need PostgreSQL or DuckLake
- opt-in live smoke tests and their exact repository-only, command-scoped fixture inputs
- expected skip behavior when live credentials are absent
- manual verification queries or file checks when no automated test exists yet
Do not make live smoke tests mandatory for normal documentation builds.
Governance And Provenance Notes
State where governance is checked and what provenance is recorded. If the workflow crosses PostgreSQL metadata and lake storage, explain the observable workflow steps and any compensating cleanup expectations instead of describing the operation as one distributed transaction.
Limitations And Follow-Up
End with known limitations and links to follow-up issues. This section is required for partial, provisioning-only, smoke-test-only, and roadmap pages.
Common limitations to call out:
- a workflow has app-layer support but no dedicated polished CLI command yet
- a workflow can use selected datastore sessions, but has no dedicated daemon workflow route yet
- language bindings are mostly roadmap thin clients over existing contracts
- live external systems require credentials and opt-in tests