One of the most common questions we hear from clients is: “Do your simulations reflect how real people will respond?” A superficially convincing answer is not enough when you’re making consequential decisions about investments, products, or communications. Clients should be able to trust that simulations tell them something useful about the people they’re trying to understand.
A few weeks ago, we preregistered predictions around the latest earnings calls from NVIDIA, Marvell, and Credo, including the questions our simulated sell-side analysts would ask. Comparing those predictions with how the calls actually unfolded lets us examine whether our simulations capture what these specific groups ask in these settings, how they frame their questions, and what makes their responses distinct. It also helps us see where our approach needs refinement and what we should test next.
Why Earnings Calls
Earnings calls give us both a public record to learn from and a real event to test against.
Analysts’ previous questions and other public material give us a basis for building their personas: the subjects they return to, the concerns they raise, and the way they ask for information.
We can then ask those simulated analysts what they would want to know at the next call. When it happens, the transcript gives us a record of what the real analysts actually asked, so we can compare our predictions with responses that did not exist when we generated them.
That comparison can go beyond whether we anticipated the right topic. Analysts hear the same company news but bring different interests and concerns. Two might ask about the same business while wanting quite different answers. We can examine whether the simulations capture those differences in topic, substance, angle, and wording.
The call also has a life beyond the questions asked in the room. It prompts written articles from journalists, investor response, market movement (or lack thereof), and a large volume of subsequent discussion. Those reactions offer further opportunities to test how information travels between audiences. Here, we focus on the first part of that picture: that is, the analysts’ questions themselves.
Why We Put Predictions on the Record
Putting those predictions on the record before the event is an important part of the test. After a call, it is easy to pick an impressive match from a large set of simulated responses. That tells us very little about how well the method performed overall.
When we put predictions on the record, we commit to the parameters of the test in advance: what information the simulation will receive, what we will ask it, how we will assess the results, and, where relevant, which simpler alternatives we will compare it with. We then run the simulation and log the responses together with the analysis method, all before the event happens.
Researchers call this preregistration. Preregistration is a familiar concept from clinical trials and psychological research: setting out the test before seeing the outcome. It reduces the scope for choosing an analysis that favours the desired result. It does not guarantee a good prediction, but it lets clients check what we committed to in advance against what actually happened.
For these three simulations, we preregistered our predicted responses on the Open Science Framework (OSF), a free research platform run by the non-profit Center for Open Science. Researchers use it to register study plans and connect them to methods and findings, so others can examine how the work was done. This supports the scientific community’s effort to make research more transparent and reproducible. In our case, OSF serves as a time-stamped record of our predictions and evaluation plans before the event.
After the earnings call happens, we can compare the recorded simulation outputs with the real transcript. Clients can see what we anticipated before we knew the answer, including where our simulated responses fall short, rather than only the examples we choose to highlight.
Testing before the event also addresses a common concern about AI: the future question cannot already have appeared in the model’s training data. The simulation still draws on an analyst’s earlier public record, just as a team preparing for the call would.
A Concrete Example from Credo
Before Credo’s Q1FY27 earnings call on 1 September 2026, we built a simulated audience of 18 analysts using public records only, including their previous questions. 13 of them subsequently asked questions on the call and were included in the evaluation.
Industry group
We asked each simulated analyst for the single question they most wanted management to answer. The question Tore Svanberg’s persona asked looked strikingly similar to his later question on the earnings call.
Before the call
Simulated Tore Svanberg
“Congratulations on another record quarter and great execution here. Bill, I was hoping to zoom in a little bit more on the optical revenue trajectory for fiscal 2027, specifically regarding how we should think about the quarterly ramp of that $600 million-plus target as we head into the second half, and what kind of gross margin mix we should expect as ZeroFlap and the PIC business scale up relative to the DSP business.”
On the call, 1 September 2026
Tore Svanberg
“Yes, thank you, and congrats on the record quarter. Bill, I was hoping you could unpack a little bit the position in optical right now. You did reiterate the $600+ million, but now you also talked about NPO and maybe even doing some system-level NPO. As we think about that $600 million, both in fiscal 2027 and fiscal 2028, how should we expect the mix to look like in terms of all the varied different components?”
As you can see, the match is not exact. Our simulated version of Svanberg asked about gross margins and the quarterly ramp; in reality, Svanberg responded to the call’s discussion of NPO and extended it into fiscal 2028. But both Svanberg and our simulation of Svanberg focused on optical revenue, referred to the $600 million target and asked about the mix of components. Even the opening and phrasing were remarkably similar. For an investor relations team preparing management for such calls, this is the level of specificity worth rehearsing against.
What the Wider Results Tell Us
As part of this exercise, we conducted three preregistered simulations. We began with NVIDIA, which explored the likely analyst agenda, management’s opening theme and investor sentiment under a specified earnings scenario. For Marvell and Credo, we asked a more specific question: does building a detailed picture of an analyst help us anticipate their questions? To test this, we compared our detailed personas with an AI model given only basic information about each analyst’s identity and role (“basic profile”). Both versions used the same underlying model, instructions, and briefing, allowing us to examine what the additional detail about the person contributed.
Three findings particularly relevant for clients:
1. Detailed personas came closer to the questions analysts actually asked.
For Marvell and Credo, their questions more closely resembled the real questions in substance and phrasing than those generated by the basic profile. For Credo, detailed personas did not improve our ability to predict the broad topics, but they did come closer to what analysts asked about those topics and how they expressed it. That’s particularly relevant when you’re preparing management for a call. While knowing a topic is likely to come up is important, anticipating the particular angle or approach an analyst might take gives you something more specific to prepare for.
2. Detailed personas produced questions that were distinctive to the analysts being modelled.
One test we ran was: could we recognise an analyst from their questions, without being told who they were? For Credo, we tested this by removing names, firms, and introductions, then asking an evaluator to match a real question to the simulated questions of one of five analysts. With a briefing containing company facts but no suggested topics for debate, the evaluator matched questions to the correct analyst more often using detailed personas than basic profiles, and more often than would be expected by chance. This suggests our personas captured some differences between individuals, beyond generating plausible questions for an analyst in general. When you’re preparing for a call, this could help anticipate what a particular analyst might ask.
3. The briefing strongly influenced what came back.
For Marvell, almost every briefed response focused on gross margins, a controversy we had named. For Credo, removing named debates from the briefing gave room for other concerns to emerge, but did not establish an improvement in topic accuracy. What this means for clients today is that the context given to personas strongly influences their simulated responses: the way an issue is framed can strongly steer what elements the simulated personas pick up — indeed, not dissimilar to how we humans work.
What This Means for Decision Makers
For a team preparing an executive for an earnings call, the broad agenda may already be clear. What takes more work is anticipating the questions beneath it: which part of the story needs explaining, which assumptions might be challenged, and where two analysts may want very different answers.
Simulations may be able to help with that preparation by capturing aspects of the questions people will later ask. A team could use those questions to develop its rehearsal, investigate a concern, or notice an angle it had overlooked. The value lies in having something specific to work through before the real conversation happens.
For market researchers, the findings help us be more precise about what we are testing. Getting the topic right, capturing the concern behind a question, and retaining differences between people are related but distinct achievements. Knowing where a simulation performs well helps researchers judge whether it is suited to the task at hand.
There is more to test. These studies involved a small number of analysts, and they do not establish how accurately a simulation measures second and third-order impacts, or the way that information spreads through the network. That requires comparisons with real human responses under different conditions.
Our answer to clients is becoming more specific as we build this evidence. We can show where a persona added information, where simpler approaches remained competitive, and how the briefing shaped the result. That gives us a firmer basis for discussing what simulations could contribute to a particular decision and what further evidence we would need for other decisions.
If you work in market intelligence, research, or strategy, we would be happy to take you through the time-stamped predictions and results, and discuss what they mean for your use case.
A Closer Look at the Research
Study design
NVIDIA was descriptive: three societies, five runs per question, each analyst’s predictions graded against the real question as a close match, partial overlap or no match. Marvell and Credo were controlled comparisons with hypotheses fixed in advance. Every arm used the same model, settings, and task. Only the information given changed: the full persona, the persona with web search, the persona unbriefed, a matched baseline with the profile replaced by name, title, employer, and coverage, a one-line prompt, and a track-record baseline of the analyst’s own past questions. Credo ran every briefed arm twice, with a briefing that named the live debates (protocol A) and with the same facts only (protocol B).
Sample
Marvell: 21 rostered analysts, 9 evaluable, 525 generated questions. Credo: 18 rostered, 13 evaluable, 840 generated questions. Five candidates per analyst per arm. Codebooks were built from 104 (Marvell) and 148 (Credo) past question turns and frozen before generation. All data used was publicly available.
Measures
Topic (M1). Jaccard overlap between the topics a blind coder assigns to the five candidates and to the analyst’s real questions.
Meaning (M2). Cosine similarity between text embeddings of each candidate and its closest real question, averaged over five.
Wording and voice (M3). Next-token predictive gain. A small fixed language model, separate from the generator, scores the real first question after a neutral preamble and after the preamble plus the five candidates . With the tokens of the real question:
in nats per token. It rewards content and phrasing together.
Identity (M4). A five-way line-up. With names and firms stripped, a judge from a different model family sees the real question and five candidates, one written for that analyst, and picks one. Three passes, majority vote, 20% chance.
Statistical checks
Paired within analysts: one-sided Wilcoxon signed-rank tests, 10,000-draw permutation checks, rank-biserial effect sizes, bootstrap intervals, Holm correction over the five primary tests per family. The line-up is also reported at analyst level, since five runs are not five people.
Limits
Removing named debates reduced the share of primary persona questions about gross margins from 97% to 20%, but did not produce an improvement in topic overlap that met the preregistered threshold. Analysts’ past questions remained a strong baseline for wording and identity. Adding their published research notes did not establish an additional benefit in the small group for which those notes were available.
Inspection and reproducibility
The registration bundles retain prompts, generated responses, evaluation rules, and analysis materials. Some proprietary inputs are withheld, with digital fingerprints recorded so the original files can be checked.