Skip to content

Knowledge base

AI concept testing: five concepts to two

By Artificial Societies
Published

AI concept testing screens product propositions before human fieldwork, so you can decide which concepts deserve further research and what needs refining. Artificial Societies uses networks of AI personas to simulate high-value audiences. In our consumer goods case study, a team narrowed five concepts to two for further iteration. We focus on the proposition and its positioning for a consequential launch decision; creative execution and proof of actual purchasing behaviour sit outside that scope.

Why use AI concept testing before human fieldwork?

A global consumer goods conglomerate needed to choose which product ideas to develop for health-conscious international consumers. Recruitment would take weeks, the audience was difficult to reach at scale, and the cost of a five-armed human experiment was prohibitive, according to our product-innovation case study. The team needed evidence while it could still change the concepts. We would use simulation at that point: to decide where further research should concentrate.

Our personas reproduce human opinion distributions with 86% accuracy against a 91% human ceiling, measured across 1,000 surveys; How accurate are AI personas? explains the measure and its limits. The benchmark concerns survey responses, so it supports a discussion about screening opinions, without establishing whether a particular product will sell. That comparison is a reason to examine how an audience is constructed before accepting its answers.

How does AI concept testing work?

For the consumer goods study, our engineers and social listening data partners gathered over 100,000 anonymised profiles, according to the case study. The material included Instagram, Facebook, TikTok, Reddit and X, alongside health forum discussions and reviews of health-adjacent products. Reviews of better-for-you foods and weight loss medications contributed observations for the audience.

We combined profiles with similar demographic and psychographic traits into lookalike consumers. We also mapped social influence networks and likely subgroup associations. The resulting society, named Health-Conscious Consumers, contained 1,498 personas segmented by health-consciousness and health-literacy, as the case study records. For a product lead, those Health-Conscious Consumers segments provide a basis for examining who understands the proposition and who values it.

Observed material is a condition of this work. We can model an audience from public signals or client-provided research where diverse observations exist. Where they do not, there is nothing to ground a persona in. Establishing that basis belongs before a discussion about which concepts to compare.

How are product concepts compared in a simulation?

The consumer goods experiment began by measuring opinion towards the product concepts. We then ran a five-armed randomised controlled trial, followed by exploration of product names and feature options, according to the case study. Each arm represented a concept in the comparison, and the design kept concepts from contaminating one another. The Health-Conscious Consumers society formed the audience for that comparison.

Stage in the consumer goods studyMaterial examinedWhat the work established
Opinion baselineThe product conceptsStarting opinions towards the concepts
Five-armed randomised controlled trialFive product conceptsA comparison without cross-contamination
Names and featuresProduct names and feature optionsFurther exploration of the product propositions

Source: Artificial Societies, consumer goods product-innovation case study.

Our concept-testing questions examine appeal and comprehension, perceived value, objections and receptive segments. These measures address different decisions. Comprehension asks whether people understand the intended proposition; perceived value asks what they make of its benefit. We would examine both before treating an appeal result as grounds to advance a concept.

Every survey question in Artificial Societies asks for a verbatim reason behind the selection by default. Our surveys ask for a verbatim reason behind each selection by default, as the case study explains. In this study, we analysed those reasons to understand what resonated and what did not. You can use that reasoning to frame a revision to a proposition, then take the revised proposition into further testing.

What results does AI concept testing produce?

The consumer goods team selected two of its five concepts for further iteration, according to our case study. The study combined quantitative results with the verbatim reasons behind selections. Its outcome was a decision about where to continue development, which is the result we would ask a screening exercise to produce.

The programme involved two rounds of quantitative and qualitative research, totalling 48 hours of fielding, according to the same case study. A round ran within 24 hours from fielding to results on an already-built society. Keep those measures separate when planning your own work: the published timing describes fielding, rather than the elapsed time needed to build the audience and complete an engagement.

The Artificial Societies case study reports almost 80% cost savings, with over twice the sample size compared with alternative human research. Its cost comparison names a human panel with half the sample size. That is an engagement-specific comparison to discuss against your proposed scope, rather than a general discount to apply to a research budget.

“Without Artificial Societies, we wouldn't have the quantitative insights that could help us narrow down product concepts effectively.”

Head of Innovation, Global Consumer Goods Conglomerate, product-innovation case study.

Can AI concept testing inform market entry?

Market-entry research asks how an unfamiliar audience will receive a proposition before an organisation commits to serving it. Surveying that audience can signal intent. Our method allows you to examine positioning through a simulated audience, provided diverse observations exist, whether public signals or your own first-party data and research. The starting requirement is evidence about the people you intend to reach.

Market size and audience reaction answer different questions. A sizing exercise quantifies the opportunity; concept and positioning work examines how people interpret the proposed offer. You might need to test how an established brand is perceived in a new market, how an unfamiliar segment understands price and value, or which objections need answering before entry. Those are questions about reception, where our simulation method applies.

For an entry decision, we would begin with the proposed positioning and the audience you expect to serve first. Compare responses by segment, then examine the objections behind them before choosing that initial audience. Our competitive-positioning research page addresses the positioning side of that decision; the consumer goods case supplies the worked example of product screening here.

What can AI concept testing validate before launch?

AI concept testing addresses whether a proposition is understood and valued, and which audience segments receive it favourably. Our scope is concept and positioning validation for consequential launch decisions. We would use simpler tools for low-stakes creative-execution choices, such as choosing an advertising treatment. The work described in the consumer goods study concerns the underlying product proposition, including names and features.

Purchasing behaviour requires a different standard of evidence. A concept test records stated preferences: what respondents say about an offer. Our published discussion of synthetic research validation explains the intention–behaviour gap, the difference between what people say and what they do. Simulation inherits that limit from survey research. A favourable response therefore gives you a proposition to investigate, without proving a purchase will follow.

We also test whether simulated answers hold together. The Artificial Societies Survey Evaluation Report, January 2026, records 89% internal coherence: agreement between questions about the same underlying attitudes. It reports a self-contradiction rate under 2%, against about 9% for human panels and up to 35% at the other end of the comparison. These measures concern consistency within survey answers. They give you further evidence to inspect when using simulated opinions to frame the next research decision; they do not measure purchases.

How should you use AI concept testing results in fieldwork?

Use the concept results to specify what human fieldwork must confirm. We recommend carrying the reasons for a selection into the next study, alongside the concept itself. In the Artificial Societies consumer goods case, verbatim explanations accompanied survey selections; that is the material we would examine when deciding which features to refine or which objection deserves a direct question.

Preserve the segment distinctions when drawing up the next research plan. Health-consciousness and health-literacy defined the Health-Conscious Consumers society, so an overall preference should lead you back to those groups before you choose the lead audience. Our recommendation is to advance a concept with an explicit account of who understood it, what they valued and what still needs testing with people. That gives the next researcher a proposition and a set of unresolved questions to investigate.

Frequently asked questions

Can unreleased product concepts stay confidential during testing?

Our consumer goods case study records that the five concepts were tested in full confidentiality and without cross-contamination. That makes it a relevant example when disclosure constrains your research. Discuss the unreleased material involved when scoping the work, alongside the audience and the decision the test needs to inform.

What happens to the observations used to build the consumer audience?

In the consumer goods study, we de-identified all behavioural traces to protect individual identities while grounding the simulation in observed behaviour. We combined profiles with similar demographic and psychographic traits into lookalike consumers. The case therefore describes an audience built through profile enrichment, rather than a set of named customers answering questions.

Can we return to an audience for later research?

Our earnings-call case study describes standing societies ready for the following quarter, with client access to rerun scenarios and retest scripts. Its four societies were Sell-Side Analysts, Buy-Side Investors, Corporate Ecosystem and Financial Media. The study includes 14 Sell-Side Analysts and 250 Buy-Side Investors, so it also gives you a worked example of distinct audience groups kept available for subsequent research. Discuss repeat use when scoping your product programme.

Does each AI persona represent a single real person?

The construction depends on the study. Our consumer goods work combined similar profiles into lookalike consumers. The transport-policy case study instead describes Washington D.C. Opinion Leaders: 1,500 personas matched one-to-one to real individuals using public data and social activity. That society spans more than 360 organisations and 800 distinct job titles. Discuss the audience you need when choosing the construction.

Where can we compare surveys with focus-group approaches?

Our audience simulation versus focus groups article covers that choice of research format. For product screening, the consumer goods example gives you a survey experiment to examine, with verbatim reasons behind the selections. Use the linked comparison when deciding how discussion-based research should fit into the wider investigation of your proposition.

Sources