A Language Model Is Not A Behavioral Model
This voice experience is generated by AI. Learn more.This voice experience is generated by AI. Learn more.Kate O’Keeffe is CEO and co-founder of Heatseeker, the AI-native platform for live-market customer validation.
gettyAt my company, we wanted to answer a simple question: Could frontier large language models predict which customer need would produce the strongest response in a live advertising experiment? To get this answer, ChatGPT, Claude and Gemini were tested against the recorded outcomes of 11 live-market experiments across four categories and international markets.
Each model received the same audience description and competing customer needs, without the market results, and had to choose one winner. ChatGPT matched the recorded winner four times, while Claude matched it three times and Gemini once.
In six of the 11 tests, all three models were wrong. In four, they selected the same wrong answer. In the clearest example of that shared error, all three selected the same option, which ranked last among eight on both click-through rate and our engagement score. All three models assumed customers wanted less effort. The strongest response came from a different practical benefit.
This small study does not establish what every model can predict, but it raises a practical question: When a model gives a plausible account of your customer, what evidence makes that account trustworthy?
The experiments compared three to nine customer needs through live advertising campaigns targeting defined audiences. Unlike survey responses about hypothetical preferences, the outcomes reflected observed campaign behavior. Our buyer engagement score aggregates those behaviors, weighting higher-intent actions such as leads more heavily than lighter engagement. It is a derived measure of engagement, not a direct measure of every customer’s motivation or a substitute for revenue.
The model exercise was retrospective and blinded. Each model received the same audience description, competing needs and option order from a completed experiment (but not its results), and was asked to select the strongest option. We then compared each model’s choice with the recorded winner for that audience.
The distinction matters. This tested whether models could identify observed winners from the information supplied, not whether they could forecast an unseen campaign with access to every execution detail. The sample was small and selected, and several tests involved demographic subgroups. Those boundaries matter when interpreting the findings.
If one business is considering an appeal built around visiting a physical location, it could resonate with one audience while falling flat with another. The business and message are the same, but the reasons for choosing may differ.
Those responses would not necessarily be contradictory descriptions of one customer. They could reflect different audiences with different needs or circumstances.
A convincing category story can flatten differences that matter commercially. A general explanation of what people want can sound reasonable. It does not tell us which people, under which circumstances or whether that appeal beats the alternatives.
Research offers related reasons for caution. Webb and Sheeran’s meta-analysis of 47 experimental tests found that changes in intention produced smaller changes in behavior. In a different setting, Bisbee and colleagues found that 48% of regression coefficients estimated from ChatGPT-generated political-opinion survey responses differed significantly from their human-survey counterparts.
Neither study establishes the limits of every current model. Both challenge the assumption that a plausible expression of preference reliably represents observed behavior.
Historical data tells us about things somebody tried. It does not contain the outcomes of every alternative nobody tested.
Imagine a business has sold “Proposition A” for five years. Now somebody proposes B. Millions of records about A may help frame the question without settling whether this audience would prefer B.
A model can generate plausible needs, sharpen competing propositions and expose assumptions. Those are useful contributions. The commercial decision still requires evidence about which alternative matters when customers encounter it.
An AI-generated customer assumption can become a slide, then a brief, then a campaign or product roadmap. Somewhere along the way, a hypothesis becomes accepted context.
Agents can accelerate that progression. One identifies a need, another develops a proposition and another creates the campaign. Every step may be logical given its inputs, even when the original assumption is wrong. The risk is not only an obviously absurd answer. It is a plausible answer passed downstream until nobody remembers that it was never tested.
Teams need a practical discipline: label customer assumptions, identify the consequential ones, test them against observable behavior and feed the findings back into the next decision. Match the evidence to the claim. Clicks demonstrate a different response from leads; leads demonstrate a different response from purchases.
The competitive advantage will not simply be how fast AI can decide. It will be how quickly a business can discover when its assumptions are wrong.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
Comments (0)
No comments yet. Be the first to share your opinion!