MaxDiff and TURF: Choosing Between Many Options (2026)

Survey Analytics
Tutorial
Updated Sep 02, 2026

Ask people to rate ten potential features on a 1-to-5 importance scale, and you'll reliably get back ten scores clustered somewhere between 3.5 and 4.5. It's not that people are being dishonest - it's that rating something "important" costs nothing, so almost everything ends up rated as at least somewhat important, and the scale stops actually separating what people would choose if they had to pick. Two established market research techniques exist specifically to get around this: MaxDiff, which forces real trade-offs to find out what genuinely rises to the top, and TURF, which solves a related but different problem - not what's most popular individually, but which combination of things reaches the most people without wasting overlap.

Table of Contents

  1. MaxDiff: Forcing Real Trade-Offs
  2. Designing a MaxDiff Study That Actually Works
  3. TURF: Finding the Right Combination
  4. How They Relate to Each Other
  5. Common Mistakes With Both Methods
  6. A Worked Example
  7. FAQ

MaxDiff: Forcing Real Trade-Offs

MaxDiff - short for maximum difference scaling, also called best-worst scaling - shows respondents a small set of items at a time, usually four to six, and asks two things at once: which one matters most to you, and which one matters least. Respondents work through several of these small sets, each containing a different combination of items, until every item has appeared often enough, alongside enough different competitors, to build a reliable picture. The scoring is simple in concept: each item's score is the percentage of times it was picked as best, minus the percentage of times it was picked as worst - an item picked as best often and worst rarely ends up with a strongly positive score, while one that's frequently picked as worst and rarely as best ends up clearly negative.

What makes this more useful than a straight rating scale is that every single choice yields two pieces of information rather than one, and - critically - there's no way to rate everything highly, since picking a "best" and a "worst" from the same small set forces a real, relative judgment every single time. This is the same underlying logic behind pairwise comparison, covered in our guide to pairwise preference testing, just extended from strict two-item duels to slightly larger sets with both a top and bottom pick each round.

Designing a MaxDiff Study That Actually Works

The quality of a MaxDiff result depends heavily on how the sets themselves are constructed, not just on collecting enough responses. Every item needs to appear roughly the same number of times across the full survey, and needs to be paired with a genuinely varied mix of other items rather than repeatedly landing next to the same few competitors - an item that happens to always appear alongside your two strongest performers will look weaker than one that happens to always appear against your two weakest, even if the two items are genuinely tied in real preference. Survey platforms built for MaxDiff handle this set construction automatically using a statistical design (commonly a balanced incomplete block design), rotating which items appear together across respondents so that every item gets a fair, comparable set of competitors over the course of the study.

Set size matters too - four items per set is the most common choice, since it's small enough to compare quickly and reliably, while still narrowing a large list down efficiently across a reasonable number of sets. Fewer than four (which collapses toward the two-item logic of pairwise comparison) loses some of MaxDiff's efficiency; more than six or seven starts to strain a respondent's ability to genuinely compare every item in the set at a glance, pushing them back toward the same satisficing shortcuts a straightforward rating scale would have invited in the first place.

TURF: Finding the Right Combination

TURF - Total Unduplicated Reach and Frequency - answers a genuinely different question: not which individual item is most popular, but which combination of a limited number of items covers the most people without wasting overlap. Picture a company that can realistically only launch three new product flavors this year, choosing from a list of ten candidates. Simply picking the three flavors that scored highest individually often backfires, because the most popular flavors tend to appeal to overlapping groups of people - launch your top three individually-ranked flavors and you might discover they all appeal to the same segment of customers, while a different combination of three, individually less popular, would have reached a meaningfully larger and more varied slice of your audience. TURF works through the possible combinations and calculates "reach" - the percentage of respondents who'd want at least one item in a given bundle - to find the combination that maximizes that number without needless redundancy between items serving the same people twice.

How They Relate to Each Other

MaxDiff and TURF are often used together, in sequence, rather than as competing choices. MaxDiff is typically run first, to get a reliable, forced-trade-off ranking of a larger list of candidates - features, flavors, messages - down to a more manageable shortlist. TURF is then run on that shortlist specifically to solve the combination problem: given that you can only ship, stock, or promote a limited number of them, which specific combination reaches the broadest, least-overlapping audience. Running MaxDiff alone tells you what's individually strongest; running TURF on top of it tells you what to actually ship together.

Common Mistakes With Both Methods

The most common MaxDiff mistake is treating the list of candidate items as smaller and more distinct than it really is - including several items that are, from a respondent's point of view, close to interchangeable (three very similar flavor variants, say) inflates how often that cluster shows up as a "best" pick relative to genuinely distinct options, simply because the cluster is competing with itself for attention across many different sets. It's worth reviewing a candidate item list specifically for near-duplicates before fielding the study, not just relying on the scoring to sort it out afterward.

The most common TURF mistake is treating reach as the only number that matters and ignoring frequency entirely - a bundle that reaches a huge share of the audience but only barely (each person picking it only weakly, or being lukewarm about it) can look identical to a bundle that reaches a smaller but more enthusiastic audience, if you're only looking at the reach percentage. Checking both reach and average frequency together - not reach alone - avoids picking a bundle that technically touches everyone but delights no one.

A Worked Example

A snack brand is deciding which three of twelve candidate flavors to launch this year. A MaxDiff exercise across all twelve narrows the field to a shortlist of six genuinely strong performers, each one clearly preferred over the weaker half of the list. Ranked individually, the top three from that shortlist are all fruit-forward flavors - and a TURF analysis reveals that those three fruit flavors overlap heavily in appeal, together reaching only 58% of respondents, since people who like one fruit flavor tend to like the others too and nobody who dislikes fruit flavors is picking up any of the three. Swapping the third fruit flavor for a savory option that individually ranked slightly lower pushes total reach to 74%, because the savory option pulls in a meaningfully different group of people the three fruit options were never going to reach on their own. The brand launches the mixed combination instead of the individually highest-scoring trio.

FAQ

Do I need special software to run MaxDiff or TURF?
Both typically require dedicated survey design and analysis tools built for the method, since MaxDiff needs a specific rotating-sets survey structure and TURF requires working through many possible combinations computationally. They're not something you'd set up in a standard rating-scale question or calculate by hand in a simple spreadsheet.

How many items can I include in a MaxDiff study?
There's no strict cap, but each respondent only sees a handful of items per set, several sets across the survey - a longer full list (20-30 items) generally needs more sets per respondent, or a larger sample, to get a reliable score for every item.

Is MaxDiff the same thing as the pairwise comparison used for testing images?
Related, not identical. Pairwise comparison is strictly two items at a time with one forced winner; MaxDiff shows slightly larger sets and asks for both a best and a worst pick each round. Both work on the same core insight - relative, forced choices beat independent ratings - applied slightly differently.

When should I reach for TURF instead of just picking the top-ranked items?
Specifically when you're limited to a small number of items you can actually ship, stock, or promote, and you're choosing from a larger list where some items likely appeal to overlapping audiences. If you're not resource-constrained, or every candidate clearly appeals to a different group already, TURF adds less value.

How many sets should each respondent see in a MaxDiff study?
Enough that every item appears a handful of times across the survey - for a list of twelve items shown in sets of four, somewhere around eight to twelve sets per respondent is a common range. Too few sets and some items won't have appeared often enough to score reliably; too many and respondent fatigue starts to undermine the quality of later picks.

Should I worry about near-duplicate items skewing a MaxDiff result?
Yes - a cluster of very similar items effectively splits its own vote across the cluster in some rounds while still crowding out genuinely distinct items in others, and reviewing your item list for near-duplicates before fielding is worth the extra care.


For more on forced-choice methods and why they beat rating scales generally, see Pairwise Comparison Testing and Beyond Averages: The Professional's Guide to Survey Analysis.

MaxDiff analysis best worst scaling TURF analysis feature prioritization survey product bundling research

Related Articles

How to Cross-Tabulate Survey Data (Without a Statistics Background) (2026

One overall number rarely tells you what's actually happening in your survey data - it's an average of groups that might be moving in completely different directions. Cross-tabulation, breaking one question's answers down by another, is where most of the real insight in a survey actually lives, and it's also where the most common analysis mistakes happen. This guide covers what cross-tabulation is, how to pick what's actually worth breaking down, why a segment can be too small to trust, and the mistakes that quietly make cross-tabs misleading instead of illuminating.

Why Your Survey Sample Might Not Represent Your Audience (2026)

A survey doesn't measure your whole audience - it measures whoever happened to respond, and those two groups are rarely identical. The people who bother to answer a survey are systematically different from the people who don't, in ways that quietly shape your results before you've analyzed a single answer. This guide covers what non-response bias actually is, how to check whether your respondents look like your real audience, and the basic idea behind weighting - correcting the imbalance after the fact, in plain terms.

Benchmarking Your Survey Results the Right Way (2026)

\"We're a 42, the industry average is 35\" sounds like a clean, reassuring comparison, right up until you look at how the industry average was actually measured and realize it was never measuring quite the same thing you were. This guide covers why cross-company benchmark comparisons are less apples-to-apples than they look, what actually makes two numbers comparable, and the benchmark that almost always matters more than any external one.

Tracking a Metric Over Time Without Mistaking Noise for a Trend (2026)

Running the same survey every quarter sounds like the simple part of survey analysis - the hard part is supposed to be the analysis itself. In practice, tracking a metric wave after wave introduces its own set of problems that a one-off survey never has to deal with: keeping the comparison genuinely apples-to-apples, accounting for seasonality, and telling a real multi-wave trend apart from a single wave that happened to wobble. This guide covers how to track a metric over time without those problems quietly undermining the comparison.

Reading Multiple-Choice Survey Results Without Getting Fooled (2026)

A select-all-that-apply question can produce a results table where every percentage adds up to well over 100%, and that's not a mistake - it's how the question works. This guide covers the specific ways multiple-choice results get misread: percentages that shouldn't be expected to sum to 100%, answer order quietly shaping which options get picked, and the difference between how many people picked something and how often it was picked overall.

We value your privacy

We use cookies and similar technologies to improve your experience, analyze site traffic, and personalize content. Learn more