Research data platform

Research data you can reproduce, share, and trust.

Upload instrument data and large files, build cohorts pinned to a point in time, analyze in your browser, and publish findings anyone can re-run. Access control, isolation, and audit are built in.

Example · Cohort pinned to a snapshotJul 2026
Live table0
cohort-7 · pinned Sep 20250

Built on open standards
  • Apache Iceberg
  • DuckDB
  • Amazon S3
  • marimo notebooks
  • nf-core pipelines
The problem

Data outlives the people who understood it.

  1. Results that cannot be re-run.The table changed, the notebook moved, and nobody recorded which version fed the figure.
  2. Data scattered across portals and drives.Instruments, imaging, and genomics each live somewhere different, with different rules.
  3. Access managed by email.Who may see what is a thread, not a policy, and it is never revoked on time.
Principles

Four commitments behind every feature.

Reproducible

Cohorts, releases, and findings pin the exact data snapshot, so a result from two years ago re-runs today.

Governed

Collection-level permissions, approval workflows, tenant isolation, and a full audit trail, checked on every request.

Open

Apache Iceberg in your own cloud account. Your data stays in open formats, with no proprietary lock-in.

End to end

Standardized data, SQL, notebooks, pipelines, and publishing in one place instead of five stitched-together tools.

Workflow

From instrument to publication.

Upload

CSV instrument data and large imaging or genomic files, with previews and viewers.

Standardize

Validate against common data elements, including the NLM CDE repository.

Analyze

SQL in the browser or on the server, notebooks, and optional pipelines.

Publish

Versioned releases, findings, and notebooks, each citing the data behind it.

Capabilities

What researchers and administrators get.

See the whole platform

Reproducibility

Reproducible cohorts

Define a cohort once and pin it to a snapshot. Re-run it any time and get identical results, even after new data arrives.

Publishing

Versioned releases and findings

Publish a release, then findings and notebooks that cite it. Each version is immutable and citable.

Data quality

Standardized instrument data

Validate CSVs against common data elements such as PHQ-9, GAD-7 and WHODAS, with about 22,000 NLM CDEs built in. Import existing NDA data.

Analysis

Query and notebooks, anywhere

SQL in the browser with DuckDB WASM, on the server, or through Athena. marimo notebooks need no environment to install.

Governance

Access control and isolation

Permission groups, access requests, curator review, and a separate bucket and catalog for each institution.

Optional module

Analysis pipelines

Launch curated pipelines such as nf-core sarek on AWS HealthOmics, with cost estimated up front and charged to the right budget.

Reproducibility

Every release records exactly what it was made from.

A release pins the data snapshot and tags it, so routine clean-up can never remove the state a published finding depends on. A reviewer can open the record and re-run the query.

Talk to us about your data
release-3 · exampleimmutable
collection
Adolescent Mood Cohort
snapshot
7f3a91c2e4b0
tag
release-3
structures
cde_phq901, cde_gad701
approved
data curator, 2026-03-12
cited by
finding-5-v2, notebook-2-v1
re-run check
identical result
Security and governance

Built for controlled-access data.

Access is decided by the application on every request, and each institution's data sits behind its own boundary.

Read the security overview

Tenant isolationA separate bucket and catalog per institution. Data does not cross the tenant wall.
Per-request authorizationPermissions are re-checked on every file fetch, so a change takes effect on the next request.
Audit trailUploads, previews, downloads, releases and permission changes are recorded.
Your cloud accountData lives in your AWS account in open table formats you can read without us.
Approval workflowsCurators approve collections and data structures before anything is released.
Quotas and rolesStorage limits per institution and per member, with tenant administrators scoped to their own tenant.
Who it's for

One platform, three ways in.

Find data and answer questions.

Browse the collections you can access, query them in the browser, and cite the exact release you used.

  • Search collections and instruments
  • SQL workbench and notebooks with no setup
  • Re-run any published finding
Contact us

Tell us about your data and your users.

We are working with a small number of institutes and consortia on early deployments. Tell us what you would want to reproduce first.