The product
Synthetic healthcare data, so you don't have to build it.
I generate synthetic patient archetypes, then produce their full record (eligibility, claims, Rx, labs, encounters, revenue) from the same people, calibrated to public benchmarks. Medicare Advantage is available now, with more lines of business and domains coming from the same engine.
The domains
Generate the person once, render every domain from them.
Each synthetic member is a full clinical life. Once the person is generated, the engine emits whatever data domain you need from that single coherent source, all joining on member_id. Eight domains ship today, with quality measures expanding on the same engine, and the rest render from the same people as the platform grows.
Eligibility & enrollment
Available nowMember-month enrollment, demographics, plan/program, and benefit status — the spine every other domain links to.
Medical claims
Available nowLine-level institutional + professional claims: diagnoses, procedures, settings, allowed/paid, and realistic adjustment chains.
Pharmacy
Available nowNCPDP-grade drug fills with refill chains, benefit phases, formulary tiers, and net-of-rebate economics.
Revenue & payment
Available nowPayer-side revenue the way a plan receives it — capitation, risk scores, and the factors behind every dollar.
Encounters
Available nowVisit-level utilization rolled up from the journey: office, ER, inpatient, SNF, outpatient, and lab encounters with length of stay, primary diagnosis, and DRG.
Labs & results
Available nowOrdered tests with result values trended to each member's conditions. The A1c tracks the diabetic, the eGFR tracks the CKD stage.
ADT feed
Available nowHL7-style admit, discharge, and register events derived from facility and ER encounters, with patient class, facility, and discharge disposition.
Member journeys
Available nowOne pseudo-chart per member: the archetype's true problem list, the HCCs actually coded this year, and a chart-note narrative. Ground truth paired with the observed record.
Quality measures
ExpandingMeasure-ready numerators, denominators, and gaps (HEDIS-style / Stars) rendered from each member's actual care.
Lines of business
Medicare Advantage today · the rest on the same engineThe product below (the files, sizes, releases, and pricing) is the Medicare Advantage line, available to download today. Every other line of business renders from the same archetype engine.
Medicare Advantage, available now
Available nowOne population, rendered eight ways.
The Medicare Advantage line renders eight linked domains: eligibility, medical claims, Rx claims, revenue, labs & results, encounters, the ADT feed, and member journeys. Each file mirrors how a plan actually receives that data stream. All eight show the same synthetic members from different angles, and they join on member_id.
Labs & results
NewResult values trend to each member's conditions. The A1c tracks the diabetic, the eGFR tracks the CKD stage, so the lab panel reads like a real chart instead of random numbers.
Member journeys (pseudo-charts)
NewOne row per member: the archetype's true problem list, the HCCs actually coded this year, and a chart-note narrative. The labeled ground-truth-plus-observed pair real data can't give you.
Eligibility
one row per member-monthThe enrollment spine. Demographics, plan/contract, dual & LIS & ESRD status, and the HCC condition flags that drive everything downstream.
- Member-month grain across 36 months
- Realistic age-ins, disenrollment, mortality
- Dual / LIS / ESRD flags that flip mid-year
- Links to every other file via member_id
Revenue (MMR)
one row per member-monthCMS payment the way a plan actually receives it. Part C & Part D capitation, V24/V28 blended risk scores, and the demographic + dual factors behind each dollar.
- V24, V28, and blended risk scores
- Part C + Part D capitation lines
- Coding-intensity + normalization applied
- Reconciles to eligibility member-months
Medical claims
one row per claim lineLine-level institutional + professional claims. DRGs, CPT/HCPCS, revenue codes, place of service, 25 diagnosis slots, allowed/paid, and realistic adjustment chains.
- Institutional + professional, multi-line claims
- Up to 25 ICD-10 diagnoses per claim
- Acute episode bundles (CHF, AMI, sepsis…)
- Adjustments, denials, reversals, paid-date lag
Rx claims (Part D)
one row per fill lineNCPDP-grade pharmacy claims. NDC-level fills with refill chains, benefit phases (deductible → coverage gap → catastrophic), formulary tiers, and net-of-rebate economics.
- Refill chains with realistic adherence
- Benefit phases + IRA Part D cap effects
- Formulary tiers, DAW codes, pharmacy NPIs
- Gross, member, plan, and net-of-rebate paid
Labs & results
one row per lab resultOrdered tests with result values trended to each member's conditions. The A1c reads high for the diabetic, the eGFR reads low for the CKD member.
- Condition-correlated result values
- LOINC codes, units, reference ranges
- Abnormal flags (H / L / N)
- Order frequency follows what's clinically indicated
Encounters
one row per encounterVisit-level utilization rolled up from the claim lines. Encounter type, length of stay, primary diagnosis, DRG, and rendering specialty, all reconciling with the claims.
- Office, ER, inpatient, SNF, outpatient, lab types
- Length of stay for facility encounters
- Primary diagnosis + DRG per encounter
- Joins to claims and ADT on the encounter
ADT feed
one row per ADT eventHL7-style admit, discharge, and register events derived from facility and ER encounters. The event stream care-management and VBC programs run on.
- A01 admit, A03 discharge, A04 register events
- Patient class, facility, admit source
- Discharge disposition (home, SNF, hospice, expired)
- Reconciles with encounters and claims
Member journeys (pseudo-charts)
one row per memberOne pseudo-chart per member: the archetype's true problem list, the HCCs actually coded this year, meds, key labs, and a chart-note narrative. The labeled truth-plus-observed pair real data can't give you.
- Ground-truth problem list vs. coded HCCs
- A chart-note narrative per member
- Utilization, medications, and key labs summarized
- The label set for risk-adjustment and coding-gap models
Referential integrity is guaranteed: every claim line, every fill, and every MMR row points back to a member that exists in eligibility for that month. You can join the full picture (diagnosis to spend to risk score to payment) without a single orphaned key.
Packaging
A free sample and a complete panel.
Packaging is simple. Start with the free 5,000-member sample, then buy the 100,000-member Complete panel when you're ready. Both ship all eight domains from the same calibrated engine, so what you test on the sample is what you get at scale.
Free sample
5,000 members · all eight domains
A 5,000-member population covering 2024–2026, full schema, full report bundle, no signup. Small enough to download in seconds and query on a laptop, big enough to judge the fidelity for yourself.
- Every one of the eight domains included
- Identical schema to the Complete panel
- Credibility audit in the report bundle
- No signup, no card
Complete panel
100,000 members · all eight domains
A 100,000-member population covering 2024–2026, with enough density to model rarer conditions, stabilize HCC prevalence, and trust tail behavior in the cost curve. One price for the whole joined population, every domain included.
- All eight domains, one flat price
- Train and validate risk-adjustment models at scale
- Benchmark a population with stable rates and tails
- Per-batch credibility audit ships with the data
All eight domains, 100,000 members
Prices are in USD. The Complete panel is the Medicare Advantage line; every other line of business renders from the same engine. See full pricing →
Versions = model family
Choose a release the way you'd choose a model.
Each version is a distinct model of reality. Newer releases capture more of the messy truth and carry a higher fidelity score. Older ones are lighter, cheaper, and still hit national control totals, with one more on the roadmap. Pick the fidelity your use case actually needs.
The releases below (v1, v2) are the Medicare Advantage line. More lines of business follow on the same versioned engine.
The AI engine: coherent comorbidity, real-HCC calibration, eight domains.
Best for: What every dataset ships as today. The highest-fidelity release, calibrated to the real CMS risk model.
Changelog · 2026-06
- 480 AI-authored patient archetypes (coherent co-occurring conditions)
- Calibrated against the real CMS HCC model end to end
- Code space widened to 55 DRGs and 199 CPT/HCPCS
- Four new domains: labs, encounters, ADT, and per-member journeys
- Per-member pseudo-charts pairing ground truth with the coded record
Persistence, seasonality, and the social-determinants MLR fix.
Best for: The prior statistical release, for cost-sensitive work or simpler dynamics.
Changelog · 2026-06
- Member-level utilization persistence (sticky high-utilizers)
- Monthly seasonality on ER / IP / surgery / office visits
- Dual MLR exceeds non-dual (corrected the v1 inversion)
- Explicit well-cohort carve-out for a realistic 5/50 spend curve
The first calibrated release. Solid control totals, simpler dynamics.
Best for: Schema validation, pipeline development, and the lineage starting point.
Changelog · 2026-05
- Four linked parquet files, full schema
- Calibrated to MA national control totals (risk, MLR, PMPM, util/1000)
- HCC-driven member journeys + acute episode bundles
- Automated credibility audit shipped with every batch
Real-claims pattern learning, SNP cohorts, provider continuity.
Best for: The next leap in fidelity: patterns learned directly from licensed real claims, not just calibrated to published aggregates.
Planned
- Pattern-learning pipeline trained on licensed real claims
- SNP cohort modeling (D-SNP / C-SNP / I-SNP)
- Provider continuity + referral networks
- Readmission + post-acute pathways to clinical targets
Fidelity scores are our composite credibility metric: the same eval, run release over release, so the trend is honest. v3 is calibrated to published benchmarks today, and the real-claims pattern learning that earns its score is the next thing on the roadmap. How we measure fidelity →
Every dataset is auditable
Each purchase ships with its own credibility audit.
You can validate every number before you trust it. Every dataset you buy (any file, any size, any release) arrives with an actuarial summary and a credibility audit generated against that exact batch. The metrics below are from the current Medicare Advantage release.
The report bundle covers the metrics an actuary or data scientist checks first: PMPM, medical loss ratio, average risk score, utilization per 1,000, HCC prevalence, and cost concentration. Each one is compared to a published CMS, MedPAC, or public-insurer benchmark, with citations. When a number drifts outside its expected band, the audit flags it openly.
Figures are illustrative. Each dataset's real audit, computed on the batch you purchase, ships in its report bundle alongside the data.
Judge the fidelity yourself.
Start with the free 5,000-member sample: all eight domains, full schema, full report bundle, no signup. Then take the Complete panel when you're ready to ship.