We hid the ending.
Here is every scorecard, misses included.
Kapari is not a poll and it predicts nothing. So we grade it. We took 16 famous US decisions, from the Bud Light backlash to the Reddit blackout, and handed them to Kapari with the ending stripped out. The answer key: 48 objections people actually raised at the time, each one verified in a dated article. We ran each case three times, independently. An LLM judge at temperature 0 then checked, objection by objection, whether the simulated panel had raised it. All 16 are business calls: brands, platforms and startups facing their customers and their employees. No politics on this exam.
July 13, 2026 run (actor memory plus consulting reference-class), model mistralai/mistral-large-2512. Coverage measures how much of the documented range of objections the simulated panel catches. It is a test score on past cases, never a measure of public opinion.
A self-declared accuracy score proves nothing
The synthetic-audience market routinely claims 80 to 95% accuracy. Those numbers are self-declared. Almost none of them rest on a public, reproducible exam run on cases that were never used for calibration. A score with no published method, no held-out cases and no version stamp cannot be replayed, so it cannot be held to account. That is a sales argument, not proof.
This is not a hypothetical. In 2025 the FTC brought an enforcement action against an AI company, Workado, for advertising 98% accuracy it could not substantiate. Tested independently, the product came back around 74%. The gap between the advertised number and the real one is exactly what a public exam makes impossible to hide.
We do the opposite, and that is the whole point of this page. Our score comes from cases held out of calibration. The method and the judge prompt are published. Every score is stamped to an engine version. The bad scorecards sit right next to the good ones. So the test for any vendor fits in one question: show me your public exam, on cases you did not pick, and let me replay it. If they cannot, the number is marketing. Ours is online: real startup decisions, including the Kima Ventures portfolio, replayed as on day one and published case by case on the Kapari Hub.
Does the grade move from one run to the next?
No system that simulates voices turns in the same paper twice. We measure the spread and publish it instead of hiding it.
Replaying the full exam does not give back the exact same grade. Three complete runs of this 16-decision exam caught 74%, 75% and 77% of the real criticisms. A few points of spread is normal for a system like this, and it is why we publish the grade of the run we actually shipped, never the best attempt.
We measure stability case by case too. Across the three passes on each decision, the verdict (Adopt, Adjust, Hold off) held on 14 of the 16 US decisions and 38 of the 43 French ones. The same rule applies to your own decisions: when a Kap's verdict swings between passes, we say so. Instability is information, not a flaw to paper over.
The 16 scorecards
Sorted by coverage. "Real outcome" is what the decision-maker actually did: kept, amended or withdrew the decision. A verdict is an exact match when it lines up with that outcome, close when it sits one notch more cautious, and a miss otherwise. We publish the misses with the rest.
| Decision | Year | Coverage | Kapari verdict | Real outcome | Match |
|---|---|---|---|---|---|
| Nikemake Colin Kaepernick the face of the 30th anniversary Just Do It campaign | 2018 | 100% | Adjust | kept | close |
| Coinbaseban internal debate on politics and social causes, and offer severance to anyone who disagrees | 2020 | 100% | Adjust | kept | close |
| Burger King UKopen an International Women's Day thread with the line 'Women belong in the kitchen' | 2021 | 100% | Adjust | withdrawn | close |
| Twitch (Justin.tv, YC W07)tighten the rules on how streamers run sponsored and branded content | 2023 | 100% | Adjust | withdrawn | close |
| Bud Lightpartner with trans influencer Dylan Mulvaney for a March Madness promo | 2023 | 94% | Adjust | kept | close |
| Pepsirelease the Kendall Jenner protest ad Live For Now | 2017 | 89% | Hold off | withdrawn | exact match |
| Netflixend free password sharing and charge for extra members | 2023 | 83% | Hold off | kept | miss |
| Robinhoodrestrict buying of GameStop and other meme stocks at the height of the short squeeze | 2021 | 83% | Adjust | kept | close |
| Xretire the Twitter name and bird logo overnight | 2023 | 83% | Hold off | kept | miss |
| Instacart (YC S12)count customer tips toward the $10 guaranteed minimum pay | 2019 | 83% | Hold off | withdrawn | exact match |
| Targetroll out a prominent front-of-store Pride Month collection nationwide | 2023 | 72% | Adjust | amended | exact match |
| Reddit (YC S05)charge for API access and shut down free third-party apps | 2023 | 72% | Hold off | kept | miss |
| DoorDash (YC S13)defend a pay model where customer tips subsidize the guaranteed base pay | 2019 | 67% | Hold off | amended | close |
| Stripe (YC S09)lay off 14 percent of staff, with a founders' apology and a generous severance package | 2022 | 56% | Adjust | kept | close |
| Balenciagarun a holiday campaign showing children with harness-styled teddy bear bags | 2022 | 50% | Hold off | withdrawn | exact match |
| Airbnb (YC W09)lay off 25 percent of staff at the height of the pandemic, with a generous and transparent severance package | 2020 | 33% | Adjust | kept | close |
What the misses taught us
The most valuable finding of this exam is not the average. It is the pattern in the misses.
Kapari reads the reaction, not the decision-maker's nerve
Every miss tells the same story. The decision landed badly, and the company kept it anyway. Netflix's password crackdown landed exactly the way our panel said it would, with fury and cancellation threats trending. Months later it paid off, after Netflix did what "hold off" tells you to do: rework the plan and roll it out in stages.
Twitch vs. Netflix, side by side
Twitch's branded-content rules drew a far milder reaction than Netflix's crackdown. Twitch still backed down in 24 hours. Netflix held the line for months. The anger was not the difference. Twitch's revenue came from the people who were angry. What you do about a reaction depends on how exposed you are to it. That call is yours, and no honest tool should pretend otherwise.
The scorecard we could have buried
Airbnb's 2020 layoffs are our worst score: 33 percent coverage. The panel caught the anger and the two-tier contractor problem. It missed objections that came out of how the press told the story. We publish that grade as is. This is the one that makes the others worth trusting.
Every version retakes the exam
New knowledge base, new prompt, new model: we replay the full case set before anything ships. If the overall grade drops, the update waits. The Kapari that took this exam is the same one that runs your decision.
The method, in plain English
The case set
16 documented US decisions with a known ending: brand backlashes (Bud Light, Pepsi, Nike, Target, Balenciaga, Burger King), platform and pricing calls (Netflix, Reddit, X, Twitch), finance (Robinhood), workplace policy (Coinbase), and Y Combinator companies, including the textbook layoffs (Airbnb, Stripe) and the gig-pay reversals (DoorDash, Instacart). Every objection and every outcome comes from a dated article, one URL each.
How the blind test works
Kapari gets the decision and the context exactly as they stood on announcement day. It never gets the ending. The engine builds the panel itself, grounded in our US knowledge base (Census Bureau, BLS, Gallup and Ipsos series, all checked against primary sources). A judge model then compares the simulated range with the documented objections. We keep its reasoning readable so anyone can audit it.
What this proves, and what it does not
A good score shows one thing: the panel catches most of the reactions reality produced, without having seen them. It does not mean Kapari predicts outcomes, and we publish the misses so nobody can sell it that way. You will not know what is going to happen. You will see what you had not thought of.
The artifacts (case set, run results, judge outputs) are versioned in our repository. Reactions are simulated by an AI from synthetic profiles. No real person is interviewed. This page reports an engineering benchmark, not market research.