Research data you can reproduce, share, and trust.
Upload instrument data and large files, build cohorts pinned to a point in time, analyze in your browser, and publish findings anyone can re-run. Access control, isolation, and audit are built in.
- Apache Iceberg
- DuckDB
- Amazon S3
- marimo notebooks
- nf-core pipelines
Data outlives the people who understood it.
- Results that cannot be re-run.The table changed, the notebook moved, and nobody recorded which version fed the figure.
- Data scattered across portals and drives.Instruments, imaging, and genomics each live somewhere different, with different rules.
- Access managed by email.Who may see what is a thread, not a policy, and it is never revoked on time.
Four commitments behind every feature.
Reproducible
Cohorts, releases, and findings pin the exact data snapshot, so a result from two years ago re-runs today.
Governed
Collection-level permissions, approval workflows, tenant isolation, and a full audit trail, checked on every request.
Open
Apache Iceberg in your own cloud account. Your data stays in open formats, with no proprietary lock-in.
End to end
Standardized data, SQL, notebooks, pipelines, and publishing in one place instead of five stitched-together tools.
From instrument to publication.
Upload
CSV instrument data and large imaging or genomic files, with previews and viewers.
Standardize
Validate against common data elements, including the NLM CDE repository.
Analyze
SQL in the browser or on the server, notebooks, and optional pipelines.
Publish
Versioned releases, findings, and notebooks, each citing the data behind it.
Reproducible cohorts
Define a cohort once and pin it to a snapshot. Re-run it any time and get identical results, even after new data arrives.
Versioned releases and findings
Publish a release, then findings and notebooks that cite it. Each version is immutable and citable.
Standardized instrument data
Validate CSVs against common data elements such as PHQ-9, GAD-7 and WHODAS, with about 22,000 NLM CDEs built in. Import existing NDA data.
Query and notebooks, anywhere
SQL in the browser with DuckDB WASM, on the server, or through Athena. marimo notebooks need no environment to install.
Access control and isolation
Permission groups, access requests, curator review, and a separate bucket and catalog for each institution.
Analysis pipelines
Launch curated pipelines such as nf-core sarek on AWS HealthOmics, with cost estimated up front and charged to the right budget.
Every release records exactly what it was made from.
A release pins the data snapshot and tags it, so routine clean-up can never remove the state a published finding depends on. A reviewer can open the record and re-run the query.
Talk to us about your data- collection
- Adolescent Mood Cohort
- snapshot
- 7f3a91c2e4b0
- tag
- release-3
- structures
- cde_phq901, cde_gad701
- approved
- data curator, 2026-03-12
- cited by
- finding-5-v2, notebook-2-v1
- re-run check
- identical result
Built for controlled-access data.
Access is decided by the application on every request, and each institution's data sits behind its own boundary.
One platform, three ways in.
Find data and answer questions.
Browse the collections you can access, query them in the browser, and cite the exact release you used.
- Search collections and instruments
- SQL workbench and notebooks with no setup
- Re-run any published finding
Own the data and control its release.
Manage members, decide what is released, and see what your collection is costing.
- Working area kept separate from published data
- Versioned releases with retraction and a recorded reason
- Storage and compute attributed to the collection
Govern a tenant without a ticket queue.
Set approval rules, permission groups and quotas for your institution.
- Permission groups and access requests
- Approval policy per tenant
- Audit history for review
Tell us about your data and your users.
We are working with a small number of institutes and consortia on early deployments. Tell us what you would want to reproduce first.