Kapari
How it works Use cases The exam Are they accurate? The Hub Pricing Partners Demo Log in Request access
Method · Panel size

How many voices does a decision panel need?

A panel earns its keep by covering ground, not by counting heads. A voice that thinks like the one before it adds volume and no information. Anyone can assert that, so we tested it on our own engine instead. Here is the protocol, the numbers it returned, and the two limits those numbers do not cover.

Short answer: around 30, and the binding constraint is not the one most people expect. We took three decisions already simulated at 100 voices. Then we drew smaller panels out of them, 400 draws at every size, and watched what changed. The verdict held at every size we tested from 10 upward. Precision improved slowly and steadily. What broke on small panels was coverage: at 10 voices, only 23% of draws contained all six stakeholder families. At 30 voices, 95% did. A panel that misses a family does not hand you a blurrier answer. It hands you a confident answer with a whole constituency missing.

iKapari assembles a panel of simulated voices. It is not a poll, not an opinion measurement, not a prediction, and no real person is interviewed.

Bigger is more reliable. That is true, and it is beside the point.

If you have ever commissioned research, your instinct says a panel of 30 is thin. That instinct is right, and it deserves a fair hearing before we take it apart.

Sampling works because there is something to estimate. A real population exists. It holds a true proportion you cannot see directly. You draw people from it at random, and the arithmetic tells you how far your draw is likely to sit from the truth. Draw more people, sit closer. The rule is sound and the math is not in dispute.

Kapari has no population behind it. It builds a set of voices deliberately, picked to clash with each other and to cover the roles and stances a decision has to survive. There is no hidden true number for those voices to approximate. Two different things, so two different rules.

If you have seen a demo running thousands of agents at once, the scale is real and it looks impressive. It buys precision on a quantity. This page is about a different question.

The margin of error, stated fairly, and why it does not apply here

Here is the rule people reach for. At 95% confidence and the worst-case split, the margin on a random sample of n people is 1.96 × sqrt(0.25/n). At n=30 that comes to 17.9 points. At n=2,000 it comes to 2.2 points. So yes: if you are estimating a proportion in a real population, 30 is nowhere near enough, and 2,000 buys you a great deal.

The formula comes with two conditions. First, a real population with a true proportion sitting inside it. Second, draws that are random and independent of one another.

Kapari fails both, and it fails them by design. Its voices are not drawn at random: they are built to clash, which is the opposite of a random draw. And they are not independent: every one of them is generated by the same model, so they share whatever that model leans toward. Run the formula on them and it returns a number. The number means nothing.

The sampling formula answers a question this product does not ask. Kapari does not estimate what share of a population holds a view. It shows a range of plausible reactions and the fault lines running through them.

So we measured it instead of arguing about it

The protocol first, because a result you cannot inspect is a claim.

Panel sizeSame verdict as the full panelDraw-to-draw variationDraws covering all 6 families
10100%±4.6 pts23%
20100%±3.1 pts79%
30100%±2.3 pts95%
50100%±1.5 pts100%
100reference0100%

Averaged across the three decisions. The table comes out of our internal saturation script, checked into our repository so the run can be replayed.

What this does not measure

Two limits, stated here rather than parked in a footnote.

This measures internal stability. It shows that a 30-voice panel reads like the 100-voice panel it came from. It says nothing about whether either one is right about the world. For a simulated panel, that check is structurally out of reach, and no amount of subsampling changes it. External validity is a separate question, and it gets a separate answer on the exam page.

The three test decisions had settled verdicts. None of them sat near a boundary. A decision balanced between two verdicts would not come with that guarantee, and we do not claim it does.

One last disclosure. These three panels ran in French, on the same engine. What the table measures is how the engine behaves as panel size falls. This is not a measurement taken on US panels.

The finding that set the floor

Below 30, the thing that breaks is not precision. Precision degrades gently, from ±4.6 points at 10 voices to ±2.3 at 30, and the verdict never flipped once across the sizes we tested. Coverage breaks hard. The stakeholder families are unequal in size, and the smallest one carries only a few voices. Draw 10 and you will usually leave it out entirely: roughly three-quarters of small draws missed a family.

A small panel does not hand you a blurrier picture. It hands you a sharp picture with a piece cut out of it. That is why 30 is the floor, and coverage is what set it, not margin.

The result we like least

Three decisions went to the same panel. All three came back at nearly the same average position, on the same verdict, with a similar spread. The model has a center of gravity, and it pulls.

More agents do not correct that. They sharpen it. Run the same thing across a much larger crowd and you get a very precise number sitting on the same average answer. That is the worse outcome, because precision reads as accuracy.

One caveat, and we would rather state it than have it read back to us. Those three decisions were three candidate versions of the same annual choice, all pointing the same way. On a varied set of real past decisions, the engine's readings spread out and land on different verdicts, so this is not a case of everything collapsing onto one answer. The pull shows up between decisions that already lean the same way.

What we built from that is the whole point of this page. If the average is where similar decisions look alike, the average is not what you came for. What separates one decision from another is where it splits, who defects, and who was never in the room. Panel diversity surfaces those. Panel volume does not.

Two design choices that came out of it

The chart we refused to build

A distribution drawn across thousands of voices looks like a readout from a real poll. It invites the reader to take it as an opinion measurement, and the reader is not wrong to: the chart is built to suggest exactly that.

We ruled it out at the design stage. Shares in Kapari stay qualified, and they stay tied to the simulated panel that produced them. They are never presented as a figure about a population.

What we spend the budget on instead

Several genuinely distinct profiles inside each stakeholder family: different stances, different roles, different generations, different levels of exposure to the decision.

A voice earns its seat only when it brings friction, a coalition, or a blind spot the panel did not already hold. The right size is reached, not decreed. You know you have hit it when adding another profile stops teaching you anything.

In the product, panels run from 24 to 100 voices, every profile editable. Thirty is where the reading stopped moving in our test, so thirty is what we recommend for a decision that matters. Go below it and you are trading coverage for speed. That is a fair trade on a quick pass and a poor one on an announcement you only get to make once.

Common questions

Is 30 a limitation dressed up as a method?

It is where the reading stopped moving, measured on the engine and published above with its protocol and its failure mode. Panels run up to 100 voices, and the 100-voice panel is the reference every smaller size is measured against.

Why not run thousands of voices anyway?

Because the extra voices buy precision on a number that has no population behind it, and because the engine's center of gravity gets sharper rather than more accurate as the crowd grows. The gain would be visual.

What is the real risk of a panel that is too small?

Missing a stakeholder family outright. At 10 voices, 23% of draws contained all six families. That is the number that set the floor at 30, not the margin.

Does any of this prove the panel is right about reality?

No, and nothing on this page claims it. Internal stability and external validity are two separate questions. The second one is graded blind on 16 past US decisions, published sheet by sheet, on the exam page.

i

Kapari assembles a panel of simulated voices and maps how those reactions break down, so you can prepare a decision before you announce it. It is not a poll. It is not an opinion measurement. It is not a prediction. No real person is interviewed, and no share shown in a result describes a population. Panel size is a design choice made to serve viewpoint diversity. It is never an argument that the panel represents anyone.

Put your next decision on the bench.

Bring a real one, the kind you only get to announce once. Read the range of reactions before the room gives you its own.

Request access → The method holds. Does it hold up against reality? See the exam, graded blind on past US decisions, every sheet published.