HomePlaybooksWhat is a quantitative concept test? Best practices
Guide · Concept testing

What is a quantitative concept test? Best practices

A quantitative concept test answers one of the most valuable questions in research: does this idea work for the people it is meant for, before real money goes into building it? Done badly, it hands you a confident but false read that can send a team in the wrong direction for months.

In brief

A quantitative concept test shows a target audience a product, feature, service or campaign idea and measures their reaction on scaled metrics, usually appeal, purchase intent, uniqueness and clarity. It exists to compare ideas and expose their weaknesses before development begins. It is not a sales forecast, and the scores are only interpretable against a benchmark and a decision rule set before the data arrives.

What a concept test is and what it is not

A concept test shows a target audience an idea and measures how appealing it is, how likely people would be to buy it, how clear it is, and how different it feels from what already exists. The aim is to compare ideas and find each one's strengths and weaknesses before committing to development.

Two things it is not, both worth being clear about at the start.

It is not a sales forecast: a high purchase-intent score does not reliably predict actual sales, because stated intent is almost always rosier than real behaviour. A survey has none of the price, competition and everyday friction of a real shop, so the scores rank ideas and expose weaknesses rather than projecting revenue.

It is not a replacement for qualitative work: the numbers tell you which concept performed better, not why. Pairing the quantitative test with a few in-depth conversations is what makes the result usable.

How to write a concept statement

The concept statement is the short description each respondent reads, and it is the foundation of the whole test. Written badly, the scores become uninterpretable, because you cannot tell whether people reacted to the idea or to the confusing way it was described.

A good concept statement is:

Clear: a target consumer who knows nothing about your internal work understands it in a single read.

Specific: it says what the product actually does, not what you hope it might one day become.

Benefit-led: it communicates the benefit to the shopper rather than the technical feature or the mechanism behind it.

Realistic: it avoids superlatives that make the idea sound implausibly good.

A dependable structure is a one-paragraph description, two or three short points on how it works, and a closing line stating the benefit. Keep the language descriptive rather than persuasive. If the copy reads like an advertisement, people react to the copywriting rather than the idea, and a strong concept in plain words can lose to a weak one dressed up in polish.

Before running the full study, test the statement on about five people. Ask them to read it and explain in their own words what the product does. If they cannot do that accurately, the statement needs rewriting rather than the concept.

How to keep concepts comparable

When testing more than one concept, every difference between them except the idea itself is a contaminant. Hold everything else steady: similar length and level of detail, the same tone and phrasing, and the same visual fidelity.

Visual fidelity is the one most often violated. If one concept appears as a finished pack render and another as rough text, the polished one wins on production value rather than merit. Where budget allows a proper visual for only one concept, use plain text for all of them.

Each concept should also carry a single idea. Bundling two benefits into one double-barrelled concept muddies the reaction, and the result will not show which part people responded to.

Choosing between monadic and sequential monadic

There are two main ways to show concepts, and the choice shapes both cost and cleanliness.

Monadic testing shows each respondent only one concept, in isolation, with different people seeing different concepts. Because nobody compares, each reading is uncontaminated, which is why monadic is considered the cleanest design. The cost is sample, since every concept needs its own group. Three concepts at 150 people each means 450 respondents in total.

Sequential monadic testing shows each respondent all the concepts, one after another, rating each before moving on. The same people evaluate everything, so the sample is far smaller and the comparison is direct. The order they are seen in introduces its own bias, which has to be managed by randomising the sequence across respondents.

The working rule is monadic when the decision is high-stakes and the concepts are very different, sequential monadic when they are close variations and budget or a head-to-head ranking matters more. Sequential monadic is the wrong choice when seeing the first concept would meaningfully change how the next is judged. Either way, keep the count modest, three to five in a sequential design, so fatigue does not erode the later ratings.

The core concept test metrics

A concept test rests on a small, consistent set of questions, each rated on a five-point scale, with the same battery applied to every concept so results stay comparable. Most scores are reported as Top 2 Box, the combined share of people choosing the two most positive options.

MetricQuestion askedWhat it tells you
AppealHow appealing do you find this?The headline read on whether the idea lands at all. Above roughly 60% Top 2 Box is a strong result in consumer categories
Purchase intentHow likely would you be to buy or use this if it were available?Almost always lower than appeal. The gap between the two is diagnostic
UniquenessHow different is this from what is already available to you?Whether the concept can create its own demand or will fight for share in a crowded category
ClarityHow clear is this to you?A check on the statement as much as the concept. Low clarity invalidates the other scores

Alongside the scales, include one or two open-ended questions. “What, if anything, would you change about this?” is the most useful, because the answers point directly at how to strengthen a concept and often explain a surprising score.

Why every score needs a benchmark

A Top 2 Box purchase intent of 48% means nothing on its own. Without a reference point, every result collapses into subjective interpretation, which is precisely what a quantitative test exists to remove.

Anchor the scores against something concrete: your current product, a previous concept's result, a competitor concept run as a control, or a norm built up across past studies. A concept scoring 61% against a current product at 47% tells a clear story. The same 61% floating on its own does not.

Setting a decision rule before fielding

Decide what pass and fail look like before the results come in. A worked example: advance a concept only if Top 2 Box purchase intent exceeds 45% and clarity exceeds 60%, and kill anything below 35% on purchase intent.

Fixing these thresholds in advance guards against the most human trap in research, quietly moving the goalposts after the fact to justify the decision you already wanted. A written, pre-set standard also makes it far easier to defend a recommendation to a stakeholder attached to a concept that simply did not test well. The discipline is borrowed from clinical research and applies here directly.

How to read concept test results

Four principles keep interpretation clean.

Do not over-read small differences: a Top 2 Box of 61% against 58% on samples of 150 sits inside the margin of error. Report the direction rather than false precision, and test for significance before declaring a winner.

Look at the whole profile rather than one number: a concept with modest purchase intent but very high uniqueness may be opening a new space rather than losing an old one. Read appeal, intent, uniqueness and clarity together.

Segment the results: an aggregate score can hide the real story. A concept may test flat overall but strongly among the core target, which is the group that actually matters.

Let the verbatims explain the numbers: when a score pattern is surprising, the open-ended answers usually explain it, which is why the two are read side by side.

A simple way to combine the two most important metrics is to read purchase intent against uniqueness.

Purchase intentUniquenessRead
HighHighA clear winner worth developing
HighLowViable but exposed to copycats
LowHighNovel but not yet compelling, may need a sharper value proposition
LowLowOne to drop

Worked example: testing three instant noodle flavours

An instant noodle brand has three new flavour concepts and wants to know which to launch. Because the three are close variations of one idea and the team wants a direct ranking, they run a sequential monadic test: 300 category buyers, each reading all three concepts one at a time in a randomised order, rating each on appeal, purchase intent, uniqueness and clarity, with a “what would you change” question after each.

Before fielding, the team sets the rule: advance any flavour with Top 2 Box purchase intent above 45%, kill anything below 30%. They load a benchmark, the brand's current best-selling flavour, as a reference point. The concepts are written to matched length and shown as plain text so no one flavour wins on a better image.

The results come back at 58% purchase intent for Flavour A, 46% for Flavour B, and 31% for Flavour C. A and B are close, and on samples this size the gap between them sits near the margin of error, so the team reports them as broadly level rather than crowning A outright.

Reading the full profile helps. Flavour B scores far higher on uniqueness, and the verbatims show people find it genuinely novel, while A is well-liked but familiar. Against the current flavour's benchmark of 44%, both clear the bar comfortably. The recommendation follows directly: develop A and B, with B positioned as the differentiated bet. It is defensible because the thresholds were set in advance.

7 common mistakes in concept testing

Writing an ad rather than a concept: persuasive copy makes people rate the writing rather than the idea. Keep the statement plain and descriptive.

Mismatched polish across concepts: different fidelity means you are testing production value. Hold every concept to the same visual standard, or use text for all.

Testing on the general population: feedback from people who would never buy the category dilutes the signal. Screen to genuine category buyers and the target profile.

Measuring liking instead of intent: whether someone likes an idea predicts market success weakly. Always include purchase or trial intent as a core metric.

Running without a benchmark: a score with nothing to compare against forces exactly the subjective judgement the test was meant to replace.

Moving the goalposts: deciding what counts as a pass only after seeing the data invites rationalising the answer you already wanted.

Ignoring data integrity: without speeder checks, attention questions and open-text screening, low-effort responses quietly corrupt the result. The data is only as good as who took the survey.

Why Flickly

Flickly runs monadic and sequential monadic concept tests against a verified consumer panel, with rotation configured from the start and response quality scored on every completed interview so low-effort responses are visible before they reach your results. Talk to us about your concept test.

Frequently asked questions

A quantitative concept test is a survey-based method that shows a target audience a product, feature or campaign idea and measures reaction on scaled metrics covering appeal, purchase intent, uniqueness and clarity, so ideas can be compared before development.
Purchase intent does not reliably predict sales, since stated intent runs higher than real behaviour when price, competition and everyday friction are absent. Scores rank concepts and expose weaknesses rather than forecasting revenue.
Top 2 Box is the combined share of respondents choosing the two most positive options on a scale, such as very and somewhat appealing on a five-point scale. It summarises a rating question in a clean and stable way.
Concept tests should carry a consistent battery of appeal, purchase intent, uniqueness and clarity on five-point scales, applied identically to every concept, plus one or two open-ended questions asking what respondents would change.
Fixing pass and fail thresholds before fielding prevents them being adjusted afterwards to justify a decision already preferred, and makes the final recommendation easier to defend to stakeholders.
Sequential monadic designs handle three to five concepts before fatigue distorts later ratings. Longer lists need either a monadic design or a screening round to identify finalists before full testing.
Concept testing measures which idea performed better without explaining why, so it complements rather than replaces qualitative work. Pairing scaled metrics with in-depth conversations gives the complete picture.

Run your next study with verified Indian consumers

Design the study, reach the right respondents, and get decision-ready insights in 72 hours.

Book a demo