MESSY INPUTSPIPELINES + CURATIONENRICHAGENTIC LOOPSYSTEM OF RECORDQUERY + ADAPTINTERFACESDO ANYTHING public studies quality violation detection + correction raw clinical data de-identified + abstracted clinical data EHR exports vendor APIs import FASTQ molecular pipelines BAM / CRAM images image-derived measurements + harmonized image stacks spatial single-cell spatial single-cell pipelines single-cell scRNA-seq pipelines flow cytometry remapping CyTOF microbiome metagenomics metagenomics pipelines post-pipeline enrichments validate fix · repeat UnifyBio adaptors Seurat · AnnData Datalog SQL Data Commons UI notebooks ML models JSON query API data science libraries Python · R · Julia · Clojure single-patient reports cohort analysis dashboards agents third-party apps your idea here public standards
71
datasets in the Pattern Data Commons
66,431
patients
76,136
samples
63,217
clinical observations and interventions
513.7M
molecular measurements

Pattern Data Commons totals as of Sept 30, 2026. Patients and samples are distinct IDs across datasets.

The idea

Collect once. Clean once. Analyze many times.

Most of the time in data science goes to hunting for errors that crept in at integration. UnifyBio moves that work to the front, where it is cheap, mechanical and reviewable.

01 / SCHEMA FIRST

Declare it, don't hard-code it

Annotate a Datomic schema with the Unify metamodel. Imports are then inferred from the schema, and the same tools work for new domains.

02 / FAIL FAST

Precise, mechanical feedback

Bad references, out of range measurement values, and gaps in data stop the import. Precise feedback lets humans and coding agents fix and rerun in seconds to minutes.

03 / IMMUTABLE

Reproduce any analysis

Every query has a time basis and an audit trail back to raw data. Rerun an old query as of any point in time.

04 / FLEXIBLE

Change the model, keep the tools

Import, query and visualization tools introspect the schema from the database. New schema doesn't require code changes.

05 / PROJECTABLE

Fit whatever comes next

The store can project data into whatever shape you need. Give each downstream tool the data it expects without 'munging'.

06 / OPEN

Built from CANDEL

Built on tooling open-sourced by the Parker Institute for Cancer Immunotherapy. Designed from day one to be schema-agnostic.

07 / 2% AI

Thoughtfully engineered by humans, for humans

Now AI provides massive leverage at the edges of the system.

The loop

Preprocess. Import. Validate. Fix. Repeat.

Each pass takes minutes and leaves an artifact you can review. Once harmonization is done, the agents can be dispensed with. Downstream analysis queries clean data instead of a pile of sed, awk and one-off scripts.

PreprocessReshape source files. Plain, deterministic code.
ImportA declarative config maps files into the schema.
ValidateReferential integrity, reference ranges and controlled vocabularies, run as Datalog.
Read the failureCauses are named and specific, so they can be fixed.
PublishA versioned database that everyone queries.
Pattern Schema

A schema that adapts to every new patient, sample, and assay it sees

The Pattern Schema (v0.3.8 shown) has 43 kinds, from subjects and timepoints to measurements and single cells. Teal kinds hold data, amber kinds are shared reference vocabularies such as genes, proteins and drugs. Click any kind to open its page in the schema browser.

PATTERN SCHEMA · v0.3.8 · open the schema browser
clinical- observation-set adverse- event clinical- observation drug nanostring- signature gene sgb dataset timepoint clinical- intervention-set treatment- regimen study-day subject assay sample clinical- intervention gdc-anatomic- site gene-product meddra- disease cell- population epitope cell-type clinical-trial- therapy measurement metabolite- feature single-cell otu pathway cnv atac-peak tcr variant genomic- coordinate drug-regimen clinical- trial comorbidity protein neo-antigen measurement- set lab-test measurement- matrix so-sequence- feature chr-acc-reg

Assays in the latest schema (v0.3.9)

mass cytometryflow cytometryPCRqPCRdPCRATAC-seqRNA-seqscRNA-seqspatial scRNA-seqDNA-seqIMPACT targeted panelFoundationOne targeted panelshotgun metagenomicsdual-column metabolomicsIFWESWGS16S rRNA sequencingTCR-seqBCR-seqLuminexspatial single-cellexpression arrayautoantibody arraySNP arrayNanoStringOlinkELISAVectraIMCMIBIIHCCODEXautoantibody targeted panel
Impact

One philosophy, from a single patient to a whole commons.

The Pattern Data Commons harmonizes patient-donated datasets, straight from the firehose of EHR and raw FASTQ, into a common schema alongside more than 50 public datasets. Unify Central serves it from a deliberately simple Clojure and Datomic monolith. Because the data is harmonized, queries stay small and the UI is server-rendered HTML.

Single-patient deep dives
Cohort analysis
Real research
Real impact
All reproducible

Multiple molecular data types and clinical sources live in one model, linked by shared identifiers and dataset context-specific identity resolution.

  • Hypermedia, not framework sprawl. Clojure, Datomic, hiccup and HTMX. The server renders, and declarative attributes drive the interactivity.
  • Every dataset is its own database. Datasets stay independent and versioned, and they are still comparable across studies.
  • Queries are plain JSON. Any language can call the query API, and the client libraries below make that easy.
  • Shims, not rewrites. Targeted projections turn the harmonized core into whatever a UI or analysis needs.
Interoperability

The same catalog, in four languages.

The patternq family shares one function catalog, the same result columns and the same plots. Work in the language your team already uses.

import patternq as pq
import patternq.dataset as pqd
import patternq.survival as pqs

# every dataset is its own database
pq.set_db(pq.resolve_db("prince-2022"))

pqd.dataset_summary()
outcomes = pqd.subject_outcomes()  # BOR / PFS / OS
pqd.variants(genes=["KRAS", "TP53"])
pqd.gene_expression(genes=["CXCL9"],
                    measurement="tpm")
Science

Harmonized data, real results.

Publications

Abstracts & posters

FALL 2018PICI and Cognitect start CANDEL.
SUMMER 2022More than 50 datasets in CANDEL.
SPRING 2023CANDEL is open-sourced.
FALL 2024UnifyBio development supported by the Rare Cancer Research Foundation, Vendekagon Labs, and Clojurists Together.
NOWUnify powers the Pattern Data Commons.
Get started

UnifyBio derives and declares the shape, checks every import mechanically, and leaves you a time-auditable store of record that people and algorithms can both query.