Every vendor in this category leads with an accuracy percentage. Kapari shows none. No approval rating, no popularity score, no support percentage, ever. That is a position, not a missing feature. Computing the number is easy. Making it honest is the part that fails, and the peer-reviewed evidence on that point is public. Here it is, along with what you get instead.
Accuracy is how this category sells. It is also the one claim nobody in it has to prove.
Most players in synthetic audiences advertise 80 to 95% accuracy. Every one of those figures is self-declared, and none has been validated by a public benchmark. The capital is real: Simile raised $100 million, Aaru more than $50 million. The proof is not. Ask a vendor for the method behind its percentage and three things are usually missing at once: the protocol, the cases held out of calibration, and the engine version the score belongs to.
A number without those three is not a measurement. It is a marketing asset. And it does not stay on the vendor's slide. It gets repeated in a board meeting by someone who did not build it, and that person now owns a claim they cannot defend. An unverifiable accuracy figure transfers risk to whoever says it out loud inside your company.
The literature does not say synthetic panels are useless. It says something narrower and much more inconvenient.
That is the takeaway. If the average is what you are buying, you cannot tell a working panel from a broken one. The average is the only thing an accuracy percentage reports.
Bisbee and colleagues compared a synthetic panel against the American National Election Studies (ANES), the standard US reference dataset. 48% of regression coefficients differed significantly, and 32% of those carried a flipped sign. Variance was strongly compressed, so the confidence intervals came out falsely narrow. The panel looked more certain than it had any right to be.
Verasight measured error that swings by topic rather than sitting at a stable rate: 12.1 points on politics, 23.4 points on health in its 2026 report. In earlier work, immigration came in at 11.3 points with the best setup. More context or a newer model sometimes made accuracy worse, not better.
Wang, Morgenstern, and Dickerson ran 3,200 human participants against 4 LLMs across 16 identities. Two distinct defects showed up: misportrayal of a group, and flattening of the diversity inside it. Marginalized groups were hit hardest. Santurkar and colleagues found the misalignment persists even when the model is explicitly steered toward a group.
Bisbee and colleagues ran the same prompt in April and again in July 2023 and got different distributions. A percentage published in the spring described a system that no longer existed by summer. Any accuracy claim without an engine version and a date is describing a machine that has since moved.
Read the boundary carefully, because it cuts both ways. None of this work concludes that simulated panels are worthless. It concludes that using one as a numeric stand-in for real people is unsafe. Argyle and colleagues wrote the founding silicon sampling paper in 2023. They showed that a model conditioned on real sociodemographic profiles can emulate response distributions close to human subgroups. That is the honest version of the promise, and it is where the idea came from. It is not evidence that the approach works in practice at the level a percentage implies.
One institution has already drawn its own line. Pew Research Center does not use AI to simulate public opinion. It cites research showing these systems stereotype groups, capture Republican views less well than Democratic ones, and understate disagreement. It still goes to real people, through probability sampling. Pew is not a distant authority here. It is the peer that built the thing this category claims to replace.
Three tiers, labeled. We do not present institute research as peer reviewed, and we do not present a press statement as a study.
Why the labels: a peer-reviewed result has passed independent review, an institute report has not, and an institutional statement is a position rather than a finding. Collapsing the three is how a category talks itself into a number. Kapari cites this work to expose the limits of its own field, its own product included, not to borrow authority from it.
We take the reliability question on its own page, with the four criteria we think you should apply to any vendor in this category: are synthetic audiences accurate?
Refusing the number is only half the argument. The other half is what replaces it, and whether you can act on it.
Start with the asymmetry, because it cuts against us too. A vendor that publishes an accuracy percentage can be contradicted by events. Kapari publishes no opinion figure, so nothing we show can be proven wrong by what happens after your announcement. We are not going to pretend that is a virtue we earned. It is a structural advantage of saying less, and you should hold us to a different kind of proof because of it. That is what the exam below is for.
What you get instead is structure. Not a scalar, but a shape you can act on.
Which voices support the decision, and the reason each one gives. Support that rests on a single argument is fragile support, and you can see that from the reasons, never from a percentage.
The undecided segments, with the specific objection holding each one back. This is the part of the panel where a change of framing still has leverage.
The voices that turn against the decision, and whether opposition stays contained or spreads to neighboring groups. Opposition that recruits is a different problem from opposition that grumbles.
The objection nobody in your room raised, surfaced by voices that have no reason to be polite. Then the wording that divides the panel least, tested against the ones that divide it most.
The engine that computes all of this is deterministic. The verdict, the weighted dispersion, the coalitions, the fault lines, and the distribution are calculated, not written by a model. The language model gives the voices their texture and nothing else. It never touches the arithmetic. The response comes back as one of five steps, from outright rejection to clear support. When that step is unstable across passes, Kapari says so instead of smoothing it away.
The panel is grounded rather than invented. Documented segments carry weights and sourced psychosocial drivers. Every reference statistic carries a source, a URL, and a verified flag. When no frame in the knowledge base matches your decision, Kapari pulls real sources at runtime, through OpenAlex and data.gov, rather than letting a model make up its own citations. And the panel is roughly thirty deliberately contrasting voices, not thousands of near-identical agents. Thirty voices that genuinely disagree tell you more about a fault line than ten thousand that were built to agree.
This is a different thing, not a smaller one. Estimating a population value and pressure-testing a decision against a range of plausible reactions are two different jobs. The research draws the boundary precisely. It rules out substitution for real people, and it leaves honest exploration standing, as long as the tool claims neither to measure, nor to predict, nor to represent. Kapari sits inside that boundary deliberately. The day we publish a support percentage, we land on the wrong side of it.
Refusing to publish a number is not refusing to be measured. Kapari publishes exactly one measured figure, and it is a different kind of figure.
We took 16 famous US decisions from the past, from the Bud Light backlash to the Reddit blackout, and submitted each one to Kapari without its ending. The answer key was 48 objections actually voiced at the time, verified one by one in dated articles. Three independent passes per case. A judge model at temperature 0 checked the results objection by objection.
Three complete runs of the 16-case exam scored 74%, 75%, and 77%. We publish the grade from the run actually deployed, not the best attempt. The run behind these figures is dated July 13, 2026 and stamped to its engine version, model mistralai/mistral-large-2512. Every sheet is public, wins and failures alike, and the exam is replayable.
Now the distinction that holds this entire page together. An exam score is not an opinion number, and it is not a claim about the future. It counts how many documented objections Kapari recovered on decisions that have already happened, with the outcome hidden from the engine. It says nothing about how many people will support your announcement next quarter, and it never will. One is a graded performance on the past, and anyone can check it line by line. The other is a figure about a population, which is the exact thing we refuse to produce. We measure ourselves. We do not measure your audience.
This matters commercially as well as intellectually. In 2025, the FTC brought an enforcement action against an AI company, Workado, for advertising 98% accuracy without substantiation. Tested independently, the product came back around 74%. An unsubstantiated accuracy claim is enforcement exposure in the United States, not a matter of manners. A vendor selling you a percentage is selling you their liability along with it.
The full scoreboard, case by case with its evidence: the US exam sheets.
Kapari assembles a panel of simulated voices and explores the structure of their reactions, so you can decide before you announce. It is not a poll, not a measure of opinion, not a prediction. No real person is interviewed, and no voice here speaks for a real individual. The research on this page is cited to expose the limits of this category, Kapari included, not to borrow its authority.
Put yours on the bench. You read the objections, the fault lines, and the blind spot before you announce, and you never get a percentage you cannot defend.