Documentation

Concepts, guides and API reference.

Start with the vocabulary. The same handful of ideas explain most of the platform.

Concepts

The vocabulary of the platform.

TenantAn institution's boundary: its people, collections and storage.
CollectionA body of research data with an owner, members and a set of data structures.
Data structureThe definition of an instrument or file type a collection expects.
Common data elementA standard question or measure, such as an item from PHQ-9, with agreed meaning and values.
Working areaWhere data is prepared. Visible to people who work on the collection.
Published areaReleased data. What readers of a collection see.
SnapshotThe exact state of a table at a moment in time. Everything reproducible is pinned to one.
CohortA saved selection of subjects, pinned to a snapshot so it can be re-run identically.
ReleaseA versioned publication of a collection's data, with its snapshot recorded.
FindingA published, versioned result that cites the release and method behind it.
NotebookA marimo notebook that runs in the browser. Notebooks can be published and versioned.
Permission groupA named set of collections granted to a set of people.
CuratorA person who reviews collections and data structures before release.
RunOne execution of an analysis pipeline, with its cost estimate, inputs and outputs.
Guides

Task-based guides.

The outline below is what we are writing for early deployments.

Planned

Getting started

Create a collection, attach a data structure, upload your first CSV.

Planned

Reproducible analysis

Pin a cohort, write a notebook, publish a finding, and re-run it.

Planned

Releasing data

Approvals, releases, retractions and the consumer view.

Planned

Administering a tenant

Members, permission groups, approval policy and quotas.

Planned

Pipelines

Launch a curated pipeline, register your own workflow, and read the outputs.

Planned

Security review pack

Architecture overview and answers to common questionnaire items.

Querying

Read any snapshot with plain SQL.

Tables are open Apache Iceberg, partitioned by collection. Passing a snapshot id reads the data as it was.

SELECT site, COUNT(DISTINCT subject_id) AS n
FROM iceberg_scan('s3:///warehouse/nda/cde_phq901/',
                  version = '')
WHERE collection_id = 12
GROUP BY site;

Illustrative. The workbench builds this for you and always filters to the collections you may read.

API

Everything the app does is an API call.

OpenAPI referenceEvery route and schema is described by OpenAPI, with an interactive reference served by your deployment.
Typed clientThe web app uses a TypeScript client generated from the same schema, so the reference and the client cannot drift.
Same rules everywhereAPI calls go through the same authorization as the web app.

Missing something?

Tell us what you need to see documented first.

Contact us