The Role of Synthetic Data in Market Research: A Thoughtful Discussion

By Josh Logan

Synthetic data has become one of the most debated topics in market research, and the appeal for business decision-makers is easy to understand. AI-generated respondents and modeled datasets promise faster answers, lower costs, greater scale, and fewer of the logistical challenges associated with traditional fieldwork.

But faster answers are not necessarily better answers, and that tension was at the center of a recent Wire webinar that I was able to attend featuring the excellent speaker and great human, Melanie Courtright, Sago’s CEO who was joined by a thoughtful and impressive guest, Maria Rydzewski of Mastercard. The conversation explored where synthetic data can create value, where it introduces risk, and what researchers need to consider before using it to guide business decisions.

Our takeaway is straightforward: synthetic data has a place in the research toolkit, but it should not be treated as an automatic substitute for human participants.

What is meant by synthetic data?‍ ‍

The term “synthetic data” is often used as if it describes a single methodology, but that is not the case. Some synthetic datasets begin with high-quality research collected from real people. Modeling is then used to extend that data, simulate scenarios, or fill carefully defined gaps.

Other approaches ask a general-purpose AI model to respond as if it were a particular type of consumer. These responses may be informed by broad internet content, legacy datasets, other synthetic outputs, or sources that are not fully visible to the researcher.

Both may be called synthetic data, but they offer very different levels of traceability and confidence. That distinction matters because research quality depends on more than the plausibility of an answer. It depends on knowing who and what the answer represents, how it was generated, and whether it is appropriate for the decision being made.

Plausibility is not the same as validity.

Generative AI is very good at producing responses that sound reasonable. That can make synthetic participants feel more reliable than they are.

‍A coherent answer, however, is not proof that it accurately reflects how a customer thinks, feels, or behaves.

Synthetic outputs may reproduce biases in their training data, smooth over meaningful differences between people, or rely on outdated assumptions. They may also struggle with emerging behaviors, niche audiences, culturally specific experiences, and decisions shaped by emotion or circumstance.

The risk is not always an obviously incorrect result. It is an answer that appears precise and credible but lacks a dependable connection to the people the research is meant to understand. As Sago’s companion article argues, confidence can grow faster than accuracy when assumptions stack and context disappears.

Where synthetic data can be useful

None of this means synthetic data should be dismissed. Used thoughtfully, it can create cost and time efficiencies.

Potential applications provided in the webinar included:

‍In these situations, synthetic data can help teams learn sooner and focus their research investment. The key is to treat the output as directional unless it has been rigorously validated for the specific use case.

Where human participants matter most

Synthetic data may help teams explore what is already known or model how familiar patterns could play out. Human participants are most valuable when the goal is to discover something new, understand why something is happening, or capture the context behind a decision.

Real people are especially important when research needs to:

  • Uncover the emotions, motivations, and cultural influences behind behavior

  • Explore new products, experiences, or categories for which little reliable data exists

  • Understand small, specialized, or rapidly changing audiences

  • Hear perspectives that may be underrepresented in existing datasets

  • Examine sensitive, personal, or complex experiences

  • Reveal tensions between what people say, what they feel, and what they do

  • Inform high-cost or difficult-to-reverse decisions

People bring lived experience to research. They can encounter an idea for the first time, respond in unexpected ways, reconsider their assumptions, and explain the circumstances shaping their choices. An experienced researcher can follow those signals, ask better questions, and uncover insights the original research design may not have anticipated.

That capacity to surprise us is not a weakness in the data. It is often where the most valuable learning begins and why we at W5 are dedicated to human-centric insights.

Synthetic data will continue to evolve, and so will the ways research teams put it to use. What are you seeing across the research landscape or within your organization? Where does synthetic data feel promising and where are you finding no substitute for real human perspectives?

Next
Next

Spotlight: Meeting Consumers Where They Are