What is MaxDiff? Understanding Maximum Difference Scaling, or Best-Worst Scaling
The complete guide on MaxDiff analysis
What is MaxDiff?
MaxDiff (Maximum Difference) Scaling, also called best-worst scaling, is a survey method in which respondents repeatedly choose the most and least important item from small sets.
MaxDiff is a discrete-choice technique developed by Jordan Louviere in 1987. The method exists because direct questioning fails at prioritisation. Asked to rate a list of features or claims, respondents call almost everything important, and the study returns no usable order. MaxDiff removes that option by making every answer a trade-off.
How does MaxDiff work?
A MaxDiff survey shows each respondent a series of small sets, typically 5 to 15 questions displaying 3 to 5 items each, and asks them to pick the best and worst item in every set. Across those sets each item appears several times against different rivals, and the accumulated pattern of choices yields a utility score per item on a single scale. The respondent flow runs as follows -
- The respondent sees a set of four or five items drawn from the full list.
- They select the item they consider most important and the item they consider least important.
- The next screen shows a different combination, and the process repeats.
- Over the full sequence, every item appears roughly three times against a varied set of competitors.
- Analysis converts the choice pattern into a score for each item in the original list.
A single screen carries more information than two selections suggest. Picking item A as best and item D as worst from a set of four establishes that A beats B, C and D, and that B and C both beat D. One screen therefore produces five paired comparisons rather than two data points.
How do you design a MaxDiff survey?
A MaxDiff design covers 8 to 25 items, shows 3 to 5 items per set, and follows the standard rule that every respondent sees each item approximately three times. That rule is what determines task count: twenty items shown four at a time, each appearing three times, requires fifteen tasks. Four design constraints have to hold simultaneously for the output to be valid.
- Randomisation: The sets each respondent sees, and the order they appear in, vary across the sample so no item is systematically advantaged by position.
- Item balance: Every item appears the same number of times across the design, so no item accumulates more exposure than another.
- Paired balance: Every item appears alongside every other item roughly equally often, so no item is repeatedly advantaged by facing weak competition.
- Item networking, also called connectivity: the sets overlap sufficiently that every item is linked to every other through chains of comparison, which is what allows all items to be placed on one scale.
Respondent fatigue is the ceiling on item count. Each additional item adds tasks at three appearances apiece, and past roughly twenty-five items the exercise becomes long enough that later choices degrade regardless of design quality. Where the list genuinely runs longer, a sparse design is the answer rather than more tasks.
Types of MaxDiff
Standard MaxDiff
Standard MaxDiff shows every respondent every item roughly three times, and suits lists of up to about twenty-five items.
Sparse or express MaxDiff
Sparse or express MaxDiff shows each respondent only a subset of a large item list, with each item appearing fewer times per person but enough times across the sample. It handles lists of fifty or more items at the cost of requiring larger sample.
Anchored MaxDiff
Anchored MaxDiff adds a question establishing whether each item is important at all, converting relative scores into a scale with a meaningful zero point and producing absolute rather than purely relative importance.
Dual-response MaxDiff
Dual-response MaxDiff asks a follow-up on whether the chosen best item would actually be acted on, which separates relative preference from genuine intent.
What sample size does a MaxDiff study need?
A MaxDiff study needs at least 300 respondents in total, and at least 200 respondents for every subgroup that will be reported separately. MaxDiff is less sample-hungry than conjoint, because each task is cognitively simpler and yields more comparisons per screen. The subgroup rule is the one most often missed. Three segments requiring separate reads means 600 respondents, not 300 split three ways. Subgroups have to be planned and quota-managed before fielding, because a segment discovered at analysis will not have an adequate base behind it.
| Analysis goal | Minimum sample | Recommended |
|---|---|---|
| Aggregate ranking only | 300 | 400 |
| Ranking plus two reportable subgroups | 400 | 600 |
| Ranking plus three or more subgroups | 200 per subgroup | 300 per subgroup |
| Individual-level utilities for segmentation | 500 | 800 or more |
| Feeding TURF or preference simulation | 500 | 800 or more |
How do you read MaxDiff results and utility scores?
MaxDiff utility scores are typically rescaled so all items sum to 100. That makes the average score equal to 100 divided by the number of items, so on a 16-item study the average is 6.25, an item scoring 12.5 is twice as important as average, and an item scoring 25 is four times average. Two estimation routes produce those scores. Counting analysis subtracts the share of respondents choosing an item as worst from the share choosing it as best, which gives a reliable aggregate ranking with no specialist software. Hierarchical Bayesian estimation models each respondent individually, producing a utility score per item per person, which is what enables segmentation, TURF analysis and preference share simulation. One interpretation caveat governs everything above. Scores are relative to the tested set only. They are not percentages of the population, and they do not establish that any item matters in absolute terms. Every item on a list could be unimportant in reality and the winner would still score highest.
MaxDiff Example: Prioritising sunscreen features
A skincare brand needs to decide which features to prioritise in its next sunscreen launch. The team has sixteen candidate features and a rating scale study that returned twelve of them above 4 out of 5, which provided no direction. The MaxDiff design shows sixteen items four at a time, each appearing three times, producing twelve tasks per respondent. Sample is 500 category buyers, quota-managed to allow separate reads on two priority cohorts. A single screen shows four features, with the respondent selecting one as most important and one as least important:
- No white cast on the skin
- Water and sweat resistance
- Contains niacinamide
- Fragrance-free formula
Results come back rescaled to sum to 100 across all sixteen items, making 6.25 the average.
| Feature | Score | Read |
|---|---|---|
| High SPF protection, SPF 50 or above | 20.8 | More than three times average, the dominant driver |
| No white cast on the skin | 15.1 | Well above average, the strongest formulation attribute |
| Lightweight, non-greasy finish | 11.6 | Above average, worth investment |
| Water and sweat resistance | 6.7 | Around average |
| Contains niacinamide | 3.0 | Half of average, deprioritise |
| Tinted formula | 1.3 | Near the bottom, not a purchase consideration |
Why does MaxDiff beat rating scales?
Rating scales produce a wall of 4s and 5s, because respondents rate nearly everything as important when nothing is given up by doing so. MaxDiff forces a trade-off inside every set, so items separate. It also produces ratio-scaled scores, meaning an item scoring 15 against another scoring 5 is three times as preferred, and it removes scale-use bias entirely. That last point matters more than it sounds. Some respondents habitually rate high and others rate low, so an average of 9 from one person and 5 from another can reflect response style rather than preference. Because MaxDiff records choices rather than numbers, individual differences in scale use cannot enter the data. This is what makes MaxDiff results comparable across segments, across markets and across cultures, where Likert results often are not.
| Attribute | Likert rating scale | MaxDiff |
|---|---|---|
| What the respondent does | Assigns a score to each item independently | Picks best and worst from a small set |
| Discrimination between items | Poor, most items cluster at the top | Strong, forced trade-offs separate items |
| Scale-use bias | Present, yea-sayers inflate all scores | Absent, no numbers are assigned |
| Scale type | Ordinal, differences not quantifiable | Ratio, differences quantifiable |
| Cross-segment comparability | Unreliable | Reliable |
| Respondent burden per item | Low | Moderate, but cognitively simpler |
MaxDiff vs Conjoint analysis: Which one should you use?
MaxDiff ranks individual items by importance. Conjoint analysis measures how people trade off attribute levels inside complete product configurations. Use MaxDiff when the question is which items on a list matter most, and conjoint when the question is what to build or what to charge. Respondent burden differs substantially. A MaxDiff task asks for two clicks against four short items, which is cognitively simple and fast. A conjoint task asks the respondent to evaluate several full product profiles, each combining multiple attributes at varying levels, which takes longer and requires more concentration. The two are complements rather than alternatives. The standard workflow runs MaxDiff first to winnow a long list of twenty or thirty features down to the five or six that matter, then feeds those winners into a conjoint exercise to model bundles and pricing.
| Attributes | MaxDiff | Conjoint analysis |
|---|---|---|
| What it measures | Relative importance of individual items on one list | How attribute levels trade off within a full choice |
| Question format | Small sets of items, pick best and worst | Complete product profiles including price, pick one |
| Respondent burden | Lighter | Heavier |
| Output | Importance score per item on a common scale | Utilities per attribute level, market simulator, price sensitivity |
| Answers the question | Which features, claims or messages matter most | What configuration should we build, and at what price |
| Design complexity | Moderate | High, requires careful attribute and level definition |
When should you use MaxDiff analysis?
MaxDiff suits any decision that requires prioritising a defined list of 8 to 25 items where a rating scale would return everything as important. Its most common applications are feature prioritisation, message testing and brand attribute ranking, and it is used across consumer goods, healthcare, financial services and technology.
- Product and roadmap prioritisation: Which features must ship in the next release and which can be deferred?
- Marketing message and claim testing: Which of twenty claims to take into creative development?
- Packaging element testing: Which elements on a pack matter most to shoppers at the shelf?
- Brand attribute prioritisation: Which associations to build and which to deprioritise?
- Drivers of customer satisfaction: Which parts of the experience carry the most weight?
- Segmentation input: Individual-level utilities clustered into groups with genuinely different preference structures.
| Business question | MaxDiff output that answers it |
|---|---|
| Which five features go into the next release? | Ranked importance scores across the full feature list |
| Which claim should lead the campaign? | Ranked claim scores, read by target cohort |
| Do our priority customers want something different? | Subgroup comparison across reportable cohorts |
| Which combination of products reaches the most people? | Individual-level utilities fed into TURF analysis |
| What should we test in a conjoint? | The shortlist of items that scored above average |
When should you not use MaxDiff?
Use conjoint analysis instead of MaxDiff when the decision requires trade-offs between feature combinations or any read on pricing. MaxDiff orders a list. It does not model how attributes interact, and it cannot estimate willingness to pay. Three further situations rule it out.
- The item list is not yet defined: MaxDiff measures the items supplied and cannot surface an item nobody thought to write. An undefined list needs qualitative work first.
- The list is very short: under about eight items, a straightforward ranking question produces the same answer in a fraction of the interview time.
- You need intensity rather than order: MaxDiff establishes that one item beats another without indicating how strongly anyone feels about either. An item ranked last may be actively disliked or simply least favourite among options people are broadly content with.
Running a prioritisation study on Flickly
Flickly runs surveys on a panel of verified consumers, using question types like ranking, rating and choice-based questions. Every completed response is scored for quality, so low-effort answers are caught before they reach your results. This matters for any prioritisation study. When a respondent rushes and answers at random, their choices look no different from a careful respondent's in the raw data. Quality scoring is what tells the two apart.
Frequently asked questions
Run your next study with verified Indian consumers
Design the study, reach the right respondents, and get decision-ready insights in 72 hours.


