Platform

Everything between the instrument and the paper.

One governed workspace to collect data, standardize it, analyze it, and publish results that anyone can re-run. Each stage keeps a record of the data it used.

1 · Collect

Bring instrument data and large files into one place.

Tabular instrument data and heavy imaging or genomic files share one collection, one set of permissions, and one audit trail.

Instrument CSVsValidated on upload and written to Apache Iceberg tables, partitioned by collection.
Large filesUploaded straight to storage, so imaging and genomic files do not pass through the application.
Previews without downloadingText, JSON, CSV, Excel, PDF, images and HTML open in the browser. Word and PowerPoint files show their content.
Genome viewerVCF files, including pipeline outputs, open in an in-browser genome viewer.
Folders that never breakMove and rename files freely. Storage keys never change, so nothing that cites a file breaks.
Working and published areasData being prepared stays in a working area. Consumers only ever see what has been released.
2 · Standardize

Measure the same thing the same way.

Data structures define what a collection expects. Attach approved structures and every upload is checked against them.

Common data elementsPHQ-9, GAD-7, WHODAS 2.0 and the DSM-5 Level 1 Cross-Cutting measure are available as standard instruments.
NLM CDE repositoryAbout 22,000 elements from the NLM CDE Repository, searchable when you define a structure.
NDA importBring existing NDA collections in, with their structures.
Curator approvalCollections and structures can require approval by a data curator. Each institution sets its own policy, and voluntary collections can skip structure approval.
3 · Analyze

Query and explore where the data lives.

Use the engine that fits the question. Every engine can read the data as it was at a chosen snapshot.

EngineRunsUse it for
DuckDB WASMIn your browserThe interactive SQL workbench, with no server round trip.
DuckDBOn the platformDefault queries, including any pinned snapshot.
Amazon Athena optionalIn your AWS accountTeams that already analyze in AWS.
Notebooks in the browsermarimo notebooks run in the browser with nothing to install. Save a figure from a notebook and embed it in a finding.
Time travelQuery a collection as of any release, cohort or finding, not just as it is today.
Permission-awareQueries are filtered to the collections you may read, whichever engine runs them.
4 · Publish

Release results with a record of how they were made.

Cohorts, releases, findings and notebooks all pin the exact snapshot they used, and the platform tags it so it cannot be cleaned away.

WorkingPrepare

Upload, validate, define cohorts. Visible to people who work on the collection.

Release v3Pin and approve

Snapshot tagged. Files move to the published area and are locked.

PublishedConsume

Readers see released data. Findings cite it and can be re-run.

Reproducible cohortsDefine a cohort once and pin it to a snapshot. Re-run it later and get identical results.
Versioned findingsFindings and notebooks are versioned, cite their data, and embed immutable figures.
Retraction with a reasonA released file cannot be quietly deleted. Retracting it is deliberate and records why.
A consumer viewThe collection page shows exactly what readers get, and only that.
Optional modules

Add compute and reach when you need it.

These are off until a platform administrator turns them on, so a deployment starts with only what it uses.

Analysis pipelines optionalRun curated pipelines such as nf-core sarek on AWS HealthOmics, or register your own Nextflow, WDL or CWL workflow. The cost is estimated before launch, results return to the collection as ordinary files, and any finished run can be launched again with the same settings.
Credits and spend limits optionalEvery collection is tied to a billing account, and members can carry spend limits. Runs reserve their maximum cost first and are charged what they actually use.
Use from your own AWS account optionalLink an IAM role in your account, such as a HealthOmics or Batch role, so it can read the files you are entitled to. The institution and the collection owner must both opt in, and access is revoked within minutes of a permission change.
Also included

The work around the work.

Role-aware dashboards

Administrators, collection owners and researchers each land on the view that matters to them.

Workspace explorer

Structures, files, cohorts, findings and releases in one tree, with a link to each.

Storage quotas

Limits per institution and per member, attributed to the collection's owner.

Audit log

Uploads, previews, downloads and permission changes recorded for review.

See it on your own data.

We are working with a small number of institutes and consortia on early deployments.

Request a demo