Skip to content

Knowledge base

Message testing and narrative testing explained

By Artificial Societies
Published

Message testing evaluates how strategic messages, stories or positioning frameworks perform with target audiences before they are publicly deployed, the definition our narrative-testing method uses; narrative testing examines strategic messages, stories and positioning frameworks before you commit to them. For a high-stakes announcement, the work compares reactions to competing options against the audience’s existing opinions. Artificial Societies uses networks of AI personas to simulate high-value audiences. We measure the change in their responses and examine their reasons, so you can revise the framing before presenting it to the people whose reactions matter.

Why use message testing before a high-stakes announcement?

Message testing has a particular use when showing someone a draft would itself have consequences. In our transport-policy case study, a global transportation leader needed to choose a narrative for policy influencers. Testing the drafts with those people would have altered the debate the company wanted to understand. Simulation let the team examine alternative positions without putting a message into that debate.

The investor relations team in our earnings-call case study faced a related constraint: selective-disclosure regulation prevented testing its prepared remarks and guidance with the intended audience. We tested unreleased figures and scripts using simulated personas. The point was to rehearse a consequential decision while the team still had scope to change what it said. The earnings-call study used four societies: Sell-Side Analysts, Buy-Side Investors, Corporate Ecosystem and Financial Media, with each persona matched one-to-one to a real individual through public data and social activity.

Our personas reproduce human opinion distributions with 86% accuracy against a 91% human ceiling, measured across 1,000 surveys; How accurate are AI personas? explains the measure and its limits. That benchmark gives you a basis for assessing our method; it does not promise that a particular announcement will produce the simulated response.

How does pre and post exposure message testing work?

Pre and post exposure message testing measures opinion before a message and again after exposure. The difference shows the shift against the audience’s baseline. We first construct personas representing the relevant stakeholder groups, establish their views on the topic, then present competing narratives and measure reactions. The narrative-testing method aims to isolate the effect of the content within the simulation.

Our crisis case study makes the sequence concrete. For a major consumer goods company, we built 1,174 AI personas across five societies. The baseline survey covered sentiment towards the company and the issue, alongside likely purchasing behaviour. The societies represented Policymakers and Regulators, Media, Influencers, Consumers, and Commercial Customers. Those purchasing responses were simulated intentions, so keep them distinct from observed purchases.

We then tested 24 messages and creatives, according to the crisis case study. Between tests, we built up and erased the societies’ memories of each message and its delivery. This addressed cross-contamination, where an earlier message affects a later response, and order effects, where the sequence changes the result. The design let us compare alternatives without carrying the preceding message into the next test.

The crisis study’s 328,576 responses pointed towards messaging that explained the company’s steps to mitigate and resolve the issue. It also identified a likely backlash against certain activations. Follow-on simulations examined delivery and where reactions differed by segment. For your announcement, we would treat that combination as useful evidence for editing both the statement and its delivery.

What should a narrative-testing report tell you?

A narrative-testing report should connect an audience’s response to a decision about the message. The narrative-testing method measures approval, emotional response, opinion shift and qualitative explanations. You can examine results for a stakeholder group as a whole, compare segments, then read individual personas’ reasons. Teneo’s study, for example, separated Washington D.C., Tech Leaders and General Population societies. We would want an explanation for a resistant segment before recommending that you adopt the highest-scoring framing overall.

In the Teneo case study, we supported Teneo in late 2025 on a confidential project for a major US company preparing a technology strategy announcement. We tested six competing narratives and identified the strongest overall option, then refined messaging for each audience segment. Teneo received a written report and an interactive enterprise platform containing approval scores, emotional sentiment and verbatim reactions. The deliverable gave the advisory team access to the responses behind the recommendation.

What does a message test change?

Message-testing results can challenge the team’s account of what an audience cares about. In our earnings-call case study, the company expected a margin reset to dominate discussion. Simulated analysts concentrated on the durability of demand and the scale of financing commitments. The published account reports that questions on the live call followed that pattern. This suggests a useful role for rehearsal: checking the questions your team is preparing to answer against the concerns the simulation raises.

The earnings-call team received a pre-call briefing, plus access to rerun scenarios and retest its script as figures firmed up. Script testing identified which version of the prepared remarks worked best with investors. The live-call comparison is a finding from that engagement, separate from the survey-distribution benchmark. The earnings-call study describes 250 Buy-Side Investors, a Corporate Ecosystem of 570 personas and Financial Media with 1,600+ personas, alongside the analyst society. Those audiences let the team examine investor response and how the story would travel through the surrounding organisations and press.

Our transport-policy case study examined a different decision: which narrative and proof points would withstand damaging coverage. After measuring baseline attitudes, we blind-tested two narratives and their supporting evidence, then introduced negative coverage and hypothetical scenarios. The work identified an existing credibility gap and showed where the arguments held or collapsed. The client received an advisory report covering positioning, proof points and risk, with access to inspect segment cuts and individual reactions.

The transport-policy study’s Washington D.C. Opinion Leaders society contained 1,500 personas spanning more than 360 organisations and 800 distinct job titles. Its 250,000 responses came from more than 170 questions. The persona and response totals are rounded figures; they describe the study’s scale, rather than a count of people interviewed.

How long does message testing take?

Published message-testing durations describe particular engagements. The table compares the work covered by each timing; the consumer-goods concept study provides a separate example of fielding time. Use the distinction when planning around an announcement date, because running a survey on an existing society and completing an engagement measure different amounts of work.

EngagementWhat was testedPublished durationWhat the duration covers
Teneo, late 2025Six strategic narrativesLess than 3 weeksWhole engagement, including society construction
Consumer goods crisis24 messages and creativesUnder 2 weeksWhole project
Transport policyTwo narratives and associated proof points3 weeksWhole engagement
Hyperscaler earnings callGuidance and alternative prepared remarks48 hoursWhole engagement
Consumer goods product conceptsFive concepts, then names and features48 hours of fielding across two roundsFielding only; under 24 hours per round on an already-built society

Sources: Artificial Societies’ Teneo, crisis, transport-policy, earnings-call and consumer-goods product innovation case studies. Teneo’s engagement ran in late 2025.

For repeat decisions, the earnings-call study describes standing societies that were ready for the following quarter. The client could revisit scenarios and scripts through its access to the platform. If you expect to return to an audience, include that requirement when discussing the work with us, alongside the announcement date and the material you need to test.

When is narrative testing the right research method?

Narrative testing fits decisions about strategic positioning: whether a proposition resonates, which arguments persuade an audience and why. Our scope is consequential launch and communications decisions. Advertising execution and landing-page optimisation sit outside that scope. The transport-policy study, with its test of proof points under damaging coverage, illustrates the kind of question we set out to answer.

The method depends on diverse observations of the audience, public or client-provided. Where those observations do not exist, there is nothing to ground a persona in. A defined audience and evidence about its members therefore belong at the start of the discussion, before you choose which messages to compare.

Our published discussion of synthetic validation also recognises the intention–behaviour gap: what people say can be a weak predictor of what they do. We use simulation to narrow the options; confirmation through human research remains a separate task where the decision requires it. For a communications team, the practical boundary is to use simulated opinion shifts to revise the announcement without treating those shifts as observed changes in purchasing or other behaviour.

Frequently asked questions

How many options should we bring to a narrative test?

Bring the competing positions you are considering, with the proof points needed to judge them. Our Teneo study tested six narratives; the transport-policy study compared two narratives with associated proof points. Those are examples of different decisions. We would use your decision to frame the comparison, rather than turn either count into a rule.

Does narrative testing need a one-to-one model of each stakeholder?

Our case studies describe different approaches to persona construction. The earnings-call study matched each persona to a real individual. The consumer-goods study combined profiles with similar demographic and psychographic traits into enriched lookalike consumers. Its inputs included anonymised profiles from Instagram, Facebook, TikTok, Reddit and X, alongside health forums and product reviews. Teneo’s study also used aggregated profiles. A proposal should explain how the personas represent the audience you need to understand.

Can we test confidential financial messages?

Our earnings-call study tested unreleased numbers and draft scripts through simulated personas. The account states that no draft, number or script left our private, secured environment. Its investor relations team could examine reactions to guidance and prepared remarks before the call, despite being unable to test them with real investors because of selective-disclosure constraints.

Can we test a product proposition before writing the message?

Our consumer-goods product innovation study compared five concepts using its Health-Conscious Consumers society of 1,498 personas in a five-armed randomised controlled trial, then explored names and features. The study narrowed the concepts to two. Survey questions also requested a verbatim reason for each choice, giving the team explanations to examine as it decided which propositions to develop further.

Can a message test help us prepare for individual analysts’ questions?

Our earnings-call study included 14 simulated sell-side analysts, representing every analyst who had asked a question on the company’s last eight calls. The team received the question each analyst was most likely to ask, in their own words. That is a concrete preparation artefact for the people delivering the remarks and answering questions afterwards.

Sources

  • Artificial Societies, method and evaluation, including the Survey Evaluation Report, January 2026. Source material verified 17 September 2026.
  • Artificial Societies, Narrative Testing and Stakeholder Reaction Simulation method pages. Source material verified 17 September 2026.
  • Artificial Societies, Product Concept Testing and Competitive Perception Testing method pages. Source material verified 17 September 2026.
  • Artificial Societies, synthetic-validation discussion, “An Imitation of an Imitation” and intention–behaviour glossary entry. Source material verified 17 September 2026.
  • Artificial Societies, Teneo case study, engagement in late 2025. Source material verified 17 September 2026.
  • Artificial Societies, crisis simulation case study. Source material verified 17 September 2026.
  • Artificial Societies, transport-policy positioning case study. Source material verified 17 September 2026.
  • Artificial Societies, hyperscaler earnings-call case study. Source material verified 17 September 2026.
  • Artificial Societies, consumer-goods product innovation case study. Source material verified 17 September 2026.