Hard-to-reach audiences: the economics of human versus synthetic research

Takeaways
- Between 25% and 39% of the audiences we have fielded describe a population where fewer than 1 in 10 people would qualify, depending on how much overlap we assume between an audience’s screening requirements. At the extreme, between 2.5% and 12.3% fall below 1%.
- No single requirement makes an audience unusual. The difficulty comes from the combination of seniority, function, employer size, industry and geography. About one in seven fielded audiences carries eight or more screening requirements, and those sit at a median estimated incidence of 4.7%.
- Audiences are defined by work rather than by person. A job-role requirement appears in 73% of them and a location requirement in 68%, against 29% for age and 16% for gender. About 65% are B2B-shaped.
- Hard to recruit is not the same as hard to model. Recruitment difficulty and synthetic modelability are separate questions, and the strongest commercial cases sit where human recruitment is expensive and synthetic accuracy can be shown.
Hard-to-reach audiences are one of the biggest economic challenges in market research. In a large sample of production audiences fielded on Evidenza, we estimate that 25% to 39% had estimated incidence below 10%, meaning relatively few people would be expected to satisfy all the requirements to qualify.
This matters because human research generally becomes harder and more expensive as qualified respondents become scarce, while synthetic research does not require recruiting another qualified person for every additional response.
That creates a potential economic advantage for synthetic research precisely where human recruitment is most difficult.
And importantly, hard to recruit is not necessarily hard to model.
Our production data suggests that highly selective audiences are common enough for this distinction to matter.
The problem is usually the combination
Consider two audiences. “UK adults 18+” is relatively straightforward to recruit. “Senior IT decision-makers at large manufacturers in the UK” is not.
No single requirement makes the second audience especially unusual. The difficulty comes from the combination: seniority, function, employer size, industry and geography.
That pattern is common in our data. About one in seven of the audiences we have fielded carries eight or more screening requirements, and those audiences have a median estimated incidence of just 4.7%.

Those medians are the central estimate, and the four groups cover all but the 2% of fielded audiences that carry no screening requirement at all. The figure rounds each median to a ratio, one in two through to one in twenty, and prints the exact percentage beside it.
Across the fielded audiences we could estimate, 25% to 39% fall below 10% estimated incidence, depending on how much overlap we assume between screening requirements. At the extreme, between 2.5% and 12.3% fall below 1%.

We use estimated incidence below 10% as a practical marker of a highly selective audience, one where relatively few people satisfy all the requirements for the study. These are estimates, not observed panel incidence. Their purpose is to tell us how selective an audience is likely to be, not its exact incidence in a human panel.
Companies often want people at work
The audience definitions help explain why highly selective populations are so common.
A job-role requirement appears in 73% of audiences in our sample. Location appears in 68%, age in 29% and gender in 16%. About 65% of the audiences are B2B-shaped.

In other words, these audiences are much more likely to be defined by what people do and where they work than by basic demographics.
They also tend to skew senior. Among audiences that specify job level, 32% of the average panel is C-suite, while only 7% sits below manager level.

Company size narrows the population further. Among audiences that specify it, 53% of the average panel works at companies with at least 1,000 employees. And 51% of audiences are single-market.
Combine role, seniority, employer size, industry and geography, and an audience that sounds fairly ordinary can become surprisingly selective.
Why selectivity changes the economics of human research
Incidence and recruitment difficulty are not the same thing. Incidence is the share of the relevant population expected to satisfy the qualifying criteria. Recruitment difficulty is the cost, time and feasibility of obtaining enough qualified human respondents at the required quality.
The two are related, but they are not identical. Modern panels can use profiling and targeted recruitment to find narrow populations. The underlying constraint remains: as qualified respondents become scarcer, obtaining enough of them generally requires more specialized recruitment, more time, higher cost, or some combination of the three.
At some point, the question can change from what will this sample cost to can we recruit enough qualified people at all.
Human research also has a particular cost structure. Each additional completed interview requires another qualified person to participate. Synthetic research does not face the same respondent-supply constraint, because generating another response does not require finding and recruiting another qualified human.
That means the relative economics of synthetic research can become more attractive as human respondents become harder to recruit. But the economic advantage depends on whether synthetic research can produce evidence good enough for the decision.
Hard to recruit is not necessarily hard to model
A population that is difficult to recruit through a human panel may still be possible to model well synthetically. These are separate questions.
Recruitment difficulty depends on how scarce qualified respondents are and how hard they are to reach. Synthetic modelability is whether the system can reproduce the decision-relevant responses of that audience accurately enough for the question and context.

Evidence coverage is one reason that may vary across audiences and questions. The same population may also be easier to model for one question than another.
So the conclusion is not that rarer audiences are automatically better for synthetic research. It is that hard-to-recruit audiences may be especially attractive where synthetic research can model them with sufficient confidence.
Where the economics get interesting
The opportunity is comparative. As an audience becomes more selective, the cost and difficulty of obtaining human evidence can rise. But that does not tell us whether synthetic research is good enough to replace or complement it.
That leads to the more important economic question: where does synthetic research gain the greatest comparative advantage as human evidence becomes harder to obtain?
The answer depends on both sides of the equation, the cost and feasibility of human recruitment, and the accuracy of synthetic research for that audience, question and decision.
So the strongest use cases are not simply the rarest audiences. They are audiences where human recruitment is difficult or expensive, synthetic research can be shown to work sufficiently well, and the remaining uncertainty is acceptable for the decision being made.
This also creates an unusual validation problem. Broad consumer populations are relatively easy to recruit, which makes them relatively easy to use as human benchmarks. Senior decision-makers, specialist professionals and narrow B2B populations are different. They may be precisely where synthetic research could save the most time and money, but they are also expensive populations against which to build human benchmarks.
The populations where synthetic research may have the greatest economic value can therefore also be the hardest ones to validate.
So the relevant comparison is not simply synthetic versus human. It is: for this audience and this decision, what can we learn, how quickly, at what cost, and with what confidence?
Our production sample tells us something important about one side of that equation. Highly selective audiences are common enough to matter. The next question is where synthetic research can answer questions about those audiences reliably enough to change the economics.
About the analysis
This analysis uses a large sample of production audience definitions from research conducted on Evidenza. The audience characteristics reported above are observed directly from those definitions.
Incidence is estimated. We do not observe actual human-panel incidence, so we calculate how selective each audience is likely to be under two different assumptions about how its screening requirements overlap. One reading treats each requirement as narrowing the population independently and gives the 39% figure. The other assumes the requirements overlap heavily and gives 25%. The purpose is to distinguish broadly accessible audiences from highly selective ones, not to estimate exact field incidence.
Every figure here is calculated on the audiences we have actually fielded rather than on drafts. Widen the denominator to every audience we could estimate and the same measure gives 31% to 47% below 10% incidence.
The sample reflects production work conducted on Evidenza and is not intended to represent the market-research industry as a whole. One organisation accounts for most of the audiences in it, so the B2B skew in particular may be substantially one client base.
Finally, this analysis makes no claim about synthetic accuracy. It tells us how selective the populations companies want to research can be. Whether we can model a particular audience well enough for a particular question and decision is a separate question, and one that has to be validated.


