Use cases
Realistic healthcare data to build on, today.
Payers, providers, value-based-care teams, health-tech companies, and actuaries all need realistic data to build against, and most of it sits behind a six-figure license or a months-long data agreement. Verisim is synthetic data calibrated to public benchmarks, so you can start the real work now instead of waiting. Medicare Advantage ships today across eligibility, claims, pharmacy, and revenue, with more coming on the same engine.
Benchmark populations and stress-test assumptions on data you can pull today.
Actuaries at payers and consulting firms, ACO and value-based-care analysts, and risk-bearing entities all run the same plays: benchmark a population, pressure-test a risk-adjustment or Stars approach, model a value-based contract, and stress-test cost and utilization assumptions. The methodology is the part you already know. The slow part is getting data you can point it at. Medicare Advantage is the panel you can pull today, with more lines of business on the way.
A national claims license runs $50k–$250k a year and takes months of legal review. That is hard to justify for exploratory analysis or a pitch you have not won, and the license terms rarely permit speculative or demo work anyway.
So benchmarks get built from stale public PDFs, hand-wavy assumptions, or a single prior engagement's data you are not really allowed to reuse. Every new question starts from a cold, dataless start.
Download a calibrated panel and stand up real benchmarks in an afternoon: risk scores, MLR, PMPM, utilization per 1,000, and dual vs. non-dual splits, all tied to published targets, with claims, eligibility, and revenue joined to the same members. Today that panel is Medicare Advantage, calibrated to CMS and MedPAC.
Build the deliverable, test the methodology end to end, and bring working numbers to the table. The data is synthetic and licensed for exactly this work, so there is nothing to clear with legal first.
Real-claims licenses are slow, expensive, and contractually limited to a named engagement, which rules out the exploratory analysis, spec work, and demos where you most need data.
Generic "synthetic" or open datasets carry their own gap: the risk scores, payment factors, and utilization curves are not calibrated to published targets, and they fall apart the moment an actuary checks them. Verisim is built to survive that check. Every release ships a credibility audit with citations.
A boutique actuarial firm is pitching a regional MA plan on a risk-adjustment audit. With the 100k-member Verisim panel, they reproduce the plan's likely V24-to-V28 risk-score shift, quantify the revenue impact, and bring a worked "here is what we would find" story to the first meeting, for the price of one panel rather than a six-figure data license.
Build, prototype, and demo your product before the data deal closes.
Whether you are a digital-health startup, a health-system innovation group, or a product team inside a payer, your software has to run on healthcare data: claims, eligibility, revenue, and soon labs and quality. The customer's data usually sits behind a DUA that takes months to close, which leaves the roadmap, the demo, and the sales motion waiting on data you do not have yet.
Getting PHI for development means HIPAA, BAAs, and a security review, which is months of work before a single engineer touches a row. Until then, the roadmap stalls and the demo is a slide.
Teams fall back on hand-built mock data, and it rarely survives contact with a clinician or an actuary: the codes do not co-occur, the costs do not add up, and the credibility of the whole demo collapses in one question.
Develop and demo on realistic, linked healthcare data with zero PHI exposure, which means no BAA, no security review, and no privacy risk. Claims, eligibility, and revenue render from the same members, so the whole product wires up to a real schema today. Medicare Advantage ships now, with more lines of business close behind.
Because the data is calibrated and clinically coherent, the demo holds up when a prospect's clinical or actuarial team pokes at it. When the customer's DUA finally closes, you point the same code at the real feed, because the schema matches.
Acquiring real PHI just to build is a multi-month compliance project (HIPAA, BAAs, vendor security review) that most teams cannot afford to start, let alone finish, before they need to ship.
Mock or toy data carries its own problem: it does not survive a clinician's scrutiny. Verisim gives you the realism of real healthcare data without the exposure, on a schema that mirrors production data, so the swap to live data is a config change rather than a rebuild.
Zero PHI, day one
Every row is synthesized and no real patient is ever represented, so there is nothing to de-identify, no BAA to sign, and no review board to clear before you start building.
Model-grade realism
Eligibility, claims, Rx, revenue, labs, encounters, ADT, and member journeys all render from one coherent set of member journeys, so a diabetic-with-CHF member stays consistent across files and across months. It is calibrated to published benchmarks, so you can build and validate risk, cost, and quality models on it, not just demo.
Same schema, real later
Build against the production schema now and swap in the customer's real data when the DUA closes. Your ingestion, models, and dashboards carry over untouched.
A care-management startup is six weeks from a design-partner demo, but the health-plan's data will not arrive for a quarter. They build their risk-stratification and outreach workflows on a Verisim panel, demo a working product on realistic members, and close the partner. When the DUA lands, they point the exact same pipeline at the real feed.
Train and backtest models against ground-truth labels.
AI/ML and analytics teams training a risk-adjustment or cost-prediction model, validating quality and HEDIS logic, or stress-testing a pipeline all hit the same walls with real healthcare data: no clean ground truth, access controls on every copy, and no way to regenerate it when you want to change one variable.
Real data has no labels you can trust. You never know a member's true risk or expected cost, only the messy realized outcome, so separating model error from data noise is guesswork.
It also cannot be shared freely across a team, and it cannot be regenerated. To test how your model behaves when coding intensity rises 5%, real data gives you no knob to turn.
Verisim panels ship with the generative ground truth behind each member: the underlying conditions and journey that produced the claims. You can validate against what the model should have found, with the realized outcome alongside it.
The dataset is reproducible and freely shareable across your whole team (no PHI, no per-seat access controls), and you can regenerate it with controlled perturbations to backtest behavior under conditions you choose.
Real healthcare data has no clean ground truth. You only see realized outcomes, so true risk and expected cost stay unobserved and your model evaluation inherits the noise.
Real data also has to move with access controls, and it cannot be regenerated with a controlled change, which puts reproducible experiments and deliberate stress tests out of reach. A generative source removes all three limits at once.
What a generative source lets you do
- Train and validate risk-adjustment and cost-prediction models against each member's underlying ground truth, with realized spend alongside it.
- Backtest a model under controlled perturbations: shift coding intensity, prevalence, or utilization and regenerate to see how predictions move.
- Validate quality and HEDIS-style logic against known numerators, denominators, and gaps rendered from each member's care.
- Stress-test pipelines and infrastructure at 1M-member, multi-file, line-level scale before you ever touch production data.
- Share one canonical, reproducible dataset across the whole team, with no DUA, no per-seat access review, and identical inputs for everyone.
- Build reproducible benchmarks and regression tests you can rerun release over release because the data regenerates deterministically.
A payer analytics team is rebuilding their cost-prediction model. They train on a 1M-member Verisim panel, evaluate against the known generative ground truth to isolate model error, then regenerate the panel with elevated coding intensity to confirm the model degrades gracefully. All of this happens before the new model sees a single real member.
Calibrated to published benchmarks today: Medicare Advantage across eligibility, medical, Rx, and revenue. Pattern-learning from licensed real data is on the roadmap. See the methodology →
Where we're going
RoadmapMedicare Advantage is the first line, with more on the same engine.
Today, Verisim Health renders Medicare Advantage across four domains (eligibility, medical, Rx, and revenue), and that is what is live right now. The engine builds a synthetic person and renders their whole data trail, so it extends to new lines of business and new domains as we add them. Here is where it is headed, labeled honestly: the lines and domains marked below are expanding, with availability noted on each.
Lines of business
Medicare Advantage today · more lines expandingMedicare Advantage
Available nowOur first line, shipping today: HCC risk adjustment, MMR revenue, Part D, dual/LIS dynamics.
Medicare FFS
ExpandingTraditional fee-for-service Medicare — the benchmark population behind most ACO and VBC work.
Commercial / employer
ExpandingWorking-age commercial populations — different age mix, benefit design, and cost curve.
Medicaid
ExpandingManaged Medicaid with its own eligibility churn, demographics, and program structure.
ACA / exchange
ExpandingMarketplace populations with HHS-HCC risk adjustment and metal-tier benefit design.
Data domains
Four live · encounters, labs & quality expandingEligibility & enrollment
Available nowMember-month enrollment, demographics, plan/program, and benefit status — the spine every other domain links to.
Medical claims
Available nowLine-level institutional + professional claims: diagnoses, procedures, settings, allowed/paid, and realistic adjustment chains.
Pharmacy
Available nowNCPDP-grade drug fills with refill chains, benefit phases, formulary tiers, and net-of-rebate economics.
Revenue & payment
Available nowPayer-side revenue the way a plan receives it — capitation, risk scores, and the factors behind every dollar.
Encounters
Available nowVisit-level utilization rolled up from the journey: office, ER, inpatient, SNF, outpatient, and lab encounters with length of stay, primary diagnosis, and DRG.
Labs & results
Available nowOrdered tests with result values trended to each member's conditions. The A1c tracks the diabetic, the eGFR tracks the CKD stage.
ADT feed
Available nowHL7-style admit, discharge, and register events derived from facility and ER encounters, with patient class, facility, and discharge disposition.
Member journeys
Available nowOne pseudo-chart per member: the archetype's true problem list, the HCCs actually coded this year, and a chart-note narrative. Ground truth paired with the observed record.
Quality measures
ExpandingMeasure-ready numerators, denominators, and gaps (HEDIS-style / Stars) rendered from each member's actual care.
We are building a synthetic-data platform for healthcare, and we would rather be precise than grand: Medicare Advantage and four domains are what ship today, and every dataset is calibrated to published benchmarks. Pattern-learning from licensed row-level data is the v3 roadmap. If a line of business, domain, or panel you need is on this list, tell us. Roadmap order follows demand.
Put it against your own use case.
Download the 1,000-member sample with the full schema and report bundle, no signup and no card. Benchmark it, demo on it, or train against it, and judge the fidelity yourself.