Methodology

    How synthetic
    research works

    Synthetic research is only worth anything if you know where it is reliable and where it is not. This page sets out how a digital respondent is built, how a test is scored, and which questions the method answers badly.

    How a digital respondent is built

    A segment is a specification, not a filter over a database. You describe who should be answering — gender, age, cities and regions, income, education, household composition, and the behaviour that matters for your category — and the platform generates a panel of personas against that specification.

    Each persona carries a full profile and holds it for the whole test. That consistency is the point: an answer to question four has to be reconcilable with the profile that produced the answer to question one, otherwise the panel is just noise with demographics attached.

    How answers are collected

    Every test type ships with a questionnaire built on the standard methodology for that kind of pretest — cut-through, comprehension, brand fit and action for static creatives; message, comprehension, likability, relevance, distinctiveness and purchase intent for video ideas; and so on across the fifteen communication metrics the platform measures.

    Questions come in two forms: scaled, answered on a 1–5 scale, and open-ended, answered in the respondent's own words. Both matter. The scale gives you a ranking; the open answers tell you why the ranking looks the way it does, and they are where a test earns its keep.

    How results are scored

    Top-2-Box
    The headline figure: the share of respondents scoring 4 or 5. It is the industry-standard measure for pretests because it counts genuine positive reaction rather than letting a mass of indifferent middle scores flatter a mean.
    Mean and Top-1-Box
    Reported alongside, because they disagree with Top-2-Box in informative ways — a high mean with a low Top-1-Box is a variant nobody dislikes and nobody loves.
    Grouped open answers
    Free-text responses are grouped by sentiment — positive, neutral, negative — and by theme, with verbatim quotes preserved so you can read the actual wording rather than a summary of it.

    How accuracy is checked

    The method is validated by parallel testing: the same material, the same questionnaire and the same segment specification are run on a digital panel and on a live audience, and the two sets of results are compared — on which variant wins, on the composition of the top three, and on which insights surface in the open answers.

    Agreement on the leading variant is the figure that matters most, because that is the decision a pretest is usually bought to make. Agreement on insights is the harder test, and the one that keeps the open-ended half of the report honest.

    Where the method does not work

    Sensory experience
    Taste, smell, texture and physical handling are outside what the model can represent. A test tells you whether a packaging idea reads well; it cannot tell you how the pack feels in a hand.
    Finished craft
    Tests are built for the idea stage. They evaluate a story, a message, a name — not the grade of a final edit or the execution quality of a finished film.
    Group dynamics
    Respondents answer independently. That removes the moderator effect and the loudest-voice effect, which is usually what you want — but it also means the method cannot show you what a real group discussion would do to an opinion.
    Genuinely novel categories
    Personas reflect established behavioural patterns. For a category that does not exist yet, and where nobody has an existing habit to reason from, treat the results as directional and validate them live.

    How to use the result

    The right place for synthetic testing is as a filter, early and often: run it where a pretest would otherwise have been skipped, use it to kill weak options before they accumulate production cost, and use the open answers as a brief for the rewrite.

    For a final decision with a large budget behind it, the strongest pattern is a pair — a synthetic test to narrow the field, live research to validate the finalist. That is cheaper than validating everything live, and considerably more defensible than validating nothing.

    Practical questions about running a test are answered in the FAQ.

    Run a free test