Pairwise Comparison (Definition, Methods, Examples, Tools)

Pairwise comparison ranks a list by showing people two options at a time and recording which one wins. It handles lists too long for drag-and-drop ranking, and it turns subjective preferences into scores you can sort. This guide covers the definition, the matrix, the three scoring methods, the sample size formula, six tools and the method’s history back to 1927.
 

What is pairwise comparison?

Pairwise comparison is the process of comparing a set of options using head-to-head pairs to judge which one is the most preferred overall. Also known as "pairwise ranking", it shows up wherever preferences need measuring: market research, strategic decisions, and voting at scale.

Pairwise Comparison Method Survey Example Pairwise Ranking Analysis Voting

This guide covers every common question about the method:

  • What is pairwise comparison?

  • How does it work?

  • Why do people use it?

  • How do you calculate the results?

  • What is a pairwise comparison matrix?

  • What are the best survey tools for it?

  • Are there different types?

  • What sample size do you need?

  • Real-life examples

  • How do you design a survey?

  • Where did the method come from?

 

How does pairwise comparison work?

The method breaks a set of options into a series of two-way votes. A participant picks one, the vote is logged, and the next pair appears.

Pairwise Comparison Calculation How To Calculate Method Formula Example

Twenty candidate names for a podcast is the kind of list that defeats people. A name has no measurable property to sort on, so ordering all twenty at once means holding nineteen comparisons in your head at the same time. Most people either freeze or answer carelessly.

Feed the same 20 names through as pairs and each decision takes a second. A few minutes later, the full order exists.

 

There are two ways to run it. One person can take on every combination, which produces an exact map of what that individual prefers. Or the work can be split, with each member of a group voting on a slice and the slices adding up to a collective ranking.

The total number of possible pairs from a list of options is n(n-1)/2, where "n" is the number of options. Ten options gives 45 possible pairs, since 10(9)/2 = 45.

Options in your list Possible pairs Practical?
5 10 Complete comparison, easily
10 45 Complete comparison, comfortably
20 190 Switch to partial
50 1,225 Partial only
100 4,950 Partial only

Further down, this guide explains the difference between a survey showing every possible pair and a partial ranking survey.

Why do people use pairwise comparison?

Easy input. Two options, one choice, no instructions needed. The decision stays quick even when the options themselves are dense or technical.

Mobile friendly. It's far easier on a phone than a drag-and-drop ranking question. Roughly 58% of survey responses now arrive from a phone (SurveyMonkey), and mobile surveys have been found to run about 10% higher completion than desktop (Kadence). A pair fits a small screen. A draggable list of 20 items doesn't.

Long lists. Every survey platform caps its drag-and-drop ranking question somewhere around 6 to 10 options, because the mental effort climbs steeply past that and people either quit or start straightlining. Pairs have no equivalent ceiling: each screen only ever asks one question, whether the list holds 12 items or 400.

Numerical results. Statements and images go in, scores come out. That conversion is what lets a subjective list survive a meeting where somebody asks for evidence.

It makes people give something up. Deciding anything means losing one option to gain another. Pairs reproduce that, where a rating scale lets people award everything an 8 and commit to nothing.

How do you calculate the results?

Voting data gets scored three ways.

1. Win rate

Wins divided by appearances, shown as a percentage or a 0 to 100 score. Ten appearances, eight wins, score of 80. This is what almost every platform uses, and the reason results stay readable by people who weren't involved in running the survey.

There's a second reading of that number worth knowing: an option scoring 80 beats a randomly chosen rival from the same list four times out of five.

OpinionX Free Tool for Pairwise Comparison Surveys

^ Example of pairwise comparison voting and results (via OpinionX). The results screen shows the Win Rate number under the “Score” column.

2. Probabilistic

Bayesian algorithms read the whole pattern of votes and estimate where each option stands against a fixed starting point, usually 1500. Each result nudges the number. ELO is the version most people have met, through chess.

Glicko and TrueSkill are the more elaborate descendants, running behind Halo, Counter-Strike and Dota. Games get away with this because players never see the raw figure. Survey stakeholders do, and 1547 isn't a number anyone can act on.

Glicko2 Glicko Algorithm Formula Calculation Method Mathematics Data Science Explained

^ Screenshot of part of the Glicko algorithm. The formula shown has since been updated to Glicko-2, which adds a variable for outcome volatility.

3. Manual

Pen and paper works under two conditions: a single voter, and a short list. The next section works through what that looks like.


Win rate carries almost all survey and research use, because the number explains itself while still holding up mathematically. Probabilistic scoring belongs to gaming, and the manual matrix mostly turns up in school math questions.


What is a pairwise comparison matrix?

A pairwise comparison matrix is a table showing your list of options along both the header row and the column, where each cell records the outcome of one head-to-head comparison.

You start on a row and compare that option against each of the columns. The row gets a 1 if it wins and a 0 if it loses.

Here's a completed matrix for four options:

Coffee Tea Juice Water Score
Coffee · 1 1 1 3
Tea 0 · 1 1 2
Juice 0 0 · 1 1
Water 0 0 0 · 0

Add up each row and sort. Coffee takes 3, Water takes none.

Where two options draw, the usual convention hands 0.5 to each side. Matrices also lean on transitivity, meaning a preference for A over B combined with a preference for B over C is taken to settle A against C without anyone voting on it.

Pairwise Comparison Matrix Manual Spreadsheet Method Calculator Example

The matrix only works for one voter and a short list. Four options needs six comparisons and fits on a page. Twenty options needs 190 and doesn't.

What are the best survey tools for pairwise comparison?

1. OpinionX

OpinionX Pairwise Comparison Tool Free Online Survey

Built for ranking surveys specifically, with teams at Disney, Google and LinkedIn among the users. The question type is called "Pair Rank", scored on win rate, with settings for forced ranking and for setting exactly how many pairs each participant votes on.

Free accounts get unlimited surveys, unlimited ranking options, both text and image formats, and as many researcher seats on the workspace as you want. One limit applies: 25 participants per survey. Lifting it costs $900 a year.

 

Vote on 10 pairs, then click the button that appears to see the overall results for everyone who's completed the survey. No login required.

What separates it from the rest of this list is analysis built for ranked results:

One-click filtering. Narrow the ranking to a single country, pricing plan or job role and watch the order change.

Segment against segment. European against American, new against tenured, defined through needs-based segmentation.

Individual rankings. Participant-level data names the people behind a score, which is how a survey becomes an interview shortlist.

All segments in one view. A contingency table covering every group at once, which is where disagreements become obvious.

How to Analyze Pairwise Comparison Results Analysis

Verdict: the only tool listed that does multiple participants, segmentation and exports in one place, and the only one where you can test all three before paying.


2. PickedShares

Mechanical engineers are the audience: PickedShares is a library of tools, frameworks and project management material with a free comparison utility inside it. Every combination appears on one page as a set of toggles.

Because it all sits on a single screen, voting is fast and the logic is obvious. It's the manual matrix with the sums handled for you. What it won't do is accept votes from anyone but you.

Verdict: solid for solo prioritisation. The surrounding page is dense with advertising, which gets in the way while you're voting.


3. PollUnit

A general-purpose poll maker with this format bundled in. The free plan stops at 20 options and 40 participants per poll; beyond that it's a monthly subscription.

The survey design is dated, unless you're a fan of animated backgrounds with shooting stars and fireflies. The company plants trees for each purchase made.

Verdict: fine for a quick team poll. Forty participants isn't a research sample.


4. AllOurIdeas

The wikisurvey concept starts here: participants don't just vote on your options, they add their own as they go. OpinionX supports the same behaviour, letting submissions join the ranking mid-survey.

The catch is maintenance. Development stopped years ago and parts have broken since. You finish with a bare ranked list, nothing to interrogate it with, and no export.

Verdict: don't pick it if analysis matters. Worth knowing about as the place wikisurveys came from.


5. 1000minds

Built in 2002 for academics and government departments wanting a bespoke decision-insights platform. The underlying format is Multi-Criteria Decision Making, a more elaborate cousin of the method in this guide. Analysis runs through a proprietary approach called PAPRIKA, which maps how variables depend on each other alongside overall priorities. Being patented, PAPRIKA isn't an algorithm you can open up and read, although the data science is documented.

No pricing appears anywhere on their site. Licences are annual and quoted individually, so a sales call stands between you and a number. There's a trial, of unstated length. Ours ran 21 days and then locked us out of our own projects.

Verdict: if you're a professional researcher doing formal multi-criteria decision analysis and you have licence budget, this beats everything else on the list, ours included. Anyone else will struggle with it without hand-holding.


6. Pairwise-Ranking-App

Single voter, like PickedShares, but the interface is closer to a real survey: one pair on screen at a time instead of a wall of toggles. Its results page is the odd one out here, reporting wins and losses per option instead of resolving them into a score.

Verdict: the tidiest solo option here if you'd rather answer one pair at a time than face a grid.

OpinionX Free Pairwise Tool

^ GIF via opinionx.co

Are there different types?

Yes. The method always follows the core principle of head-to-head voting, and there are five ways to customise a survey. These formats aren't mutually exclusive and can be combined.

Complete pairwise comparison

Nobody skips anything: every participant votes on the full set, so what comes back has no gaps in it. Run the formula on 15 options and you're asking for 105 votes.

Three situations justify that ask. The participant pool is tiny, one person is ranking their own list, or the list stops somewhere around 15 to 20 options.

Complete Pairwise Comparison vs Partial Pairwise Comparison Sample Size Robustness Confidence Number of Votes

^ Configuring the number of votes on a pairwise comparison survey with 10 options (via opinionx.co)

Partial pairwise comparison

Sampling replaces completeness once the group grows or the list passes 20 options. Each person votes on a slice, and the slices add up.

Coverage is what matters, not individual completeness. Aim for every possible pair to be seen three times somewhere in the dataset, which is 3n(n-1)/2. Divide by however many participants you're confident will finish, and that's each person's workload. Most OpinionX surveys end up here instead of on complete.

This calculator is built into every OpinionX survey that includes a Pair Rank question.

Forced pairwise comparison

Most surveys offer a third button beside the two options: skip. It exists so nobody is forced to choose between two options they find irrelevant or genuinely incomparable. Take it away and every pair must be answered, which is forced ranking.

Forced Pairwise Comparison Ranking Survey Setup Tool OpinionX

Image comparison

Text isn't a requirement. Options can be images or GIFs, which is what makes image ranking the standard approach to concept testing, where what's being judged is visual.

Comparative Preference Testing Example Survey

Adaptive comparison

Here the survey watches your earlier answers and picks the next pair accordingly. Transitivity does part of the work: pick A>B then B>C and it infers A>C, so that pair never appears. The other input is the dataset as a whole, keeping coverage even across options or steering votes to the ones that need them.

What sample size do you need for a pairwise comparison survey?

Show every pair to every participant and the question disappears. A complete comparison is already the soundest picture of one person's preferences available.

Three things usually rule that out. A 30-option list demands 435 votes, which is more than anyone will sit through. A large pool makes completeness unnecessary, because sampling each person reconstructs the aggregate anyway. And unpaid participants abandon long surveys.

The working formula is 3n(n-1)/2. Every pair should appear three times somewhere in the survey. Take that total, divide by the smallest number of participants you realistically expect, and the answer is how many votes to assign each person.

Below 10 votes each there's little point capping it, since 10 pairs is under a minute of work. Where the formula matters is when it hands back something like 34 and you need to decide whether people will actually finish.

It also runs backwards. Give it a list length and it returns the number of participants to recruit; give it a participant count and it returns the longest list you can safely use.

Examples of pairwise comparison in real-life scenarios

Ranking everything in the world

7,188 options. 25,830,078 theoretical combinations. 1.2 million actual votes.

Tom Scott's 2020 attempt to rank everything in the world is the largest partial comparison most people have come across, and it works precisely because completeness was never required. Win rate tolerates sampling. Pizza, sleep and gravity all made the top 10.

 

Assumption testing at Labster

Labster's sales team knew which feature customers wanted. Tudor Cristian Bogdan, a UX researcher there, wasn't convinced. Instead of arguing, he pulled recent customer feedback together through internal workshops and put the whole lot through a pairwise survey. The certain feature missed the top 10, and the team's approach to internal assumptions changed after that.

Message testing at OpinionX

Our own launch went badly. 150+ interviews in and nobody was converting.

So we took every problem those interviews had raised, loaded them into a pairwise survey, and had the answer inside two hours: the problem we'd built the company around finished last. Five problems above it were ones the product already handled. New messaging went up and paying customers arrived that week. Customer Problem Stack Ranking has the full account.

 

Roadmap prioritisation at Safe.Global

Safe.Global builds wallet and account infrastructure on Ethereum, in an industry where users expect a say in the tools they depend on. Every quarter, the team puts its roadmap to the community as a pairwise vote and lets the ranking decide what comes next.

We knew our old process wasn’t a very solid approach to say which things should be prioritized.
— Kristina Mayman, UX Researcher at Safe.Global

What changed for her team was being able to sort the whole roadmap by what users said mattered most, instead of arguing it out internally.

The Wikimedia Foundation product roadmap

The Wikimedia Foundation put its 2014 tooling roadmap to the community as a prioritisation survey. More than 30,000 votes came back.

Wikimedia Wikipedia Pairwise Comparison Survey 2014 Community Survey

Player rating on Chess.com

A chess game is already a pair vote with players instead of options. Chess.com, the largest chess community anywhere, runs Glicko-2 across millions of them, seeding every account at 1500 and moving the number with every result.

Idea validation at Stripe

Validation methods usually stay inside the company that built them. Shreyas Doshi, then a product leader at Stripe, published his in a post called "Destined To Fail", setting out why most idea validation fails. His remedy was to rank customer problems comparatively, because the position tells you whether the problem you picked was ever near the top of theirs. Full story here.

Academia and formal research

The academic use is the old one; product teams and executives only picked it up recently. Both now sit on the same tooling. Disney, Google, LinkedIn, Shopify and Amazon run internal and customer-facing surveys with it, while academic users apply it to social impact, educational engagement and medical research.

How do you design one?

Setting one up takes no more work than any other survey type. Two requirements, two suggestions.

1. The comparison question

What lens should participants use to interpret the pair they're voting on? The six most common:

  • Preference. The plain one. Which appeals more?

  • Pain. Where does the unmet need bite harder?

  • Value. Which is worth more money to them?

  • Risk. Which one worries them?

  • Motivation. Which is likelier to make them act?

  • Friction. Which is likelier to stop them?

A lens alone isn't a question. It needs a context to sit in. Combine the value lens with your product's paid feature set and you land on something answerable: "Which of these two features is worth more of what you pay us?"

2. The comparison options

The list is what gets voted on. Sticking with the value question above, that means every feature currently shipping, not the ones on the roadmap.

In user research the list is usually problem statements instead of features, run under the name Customer Problem Stack Ranking. Either way, participants can add to the list mid-survey if you'd rather crowdsource it than write it all yourself.

3. Participant identifiers

Anonymity is the default here, which is the wrong setting for most research. An identifier question captures names, emails or usernames against each set of votes.

Skip it and you lose the follow-up. The moment a result surprises you, the next question is who voted that way, and identifiers are what turn a ranking into a list of people to talk to. The Discovery Sandwich is the method that depends on it.

4. Segmentation data

A single aggregate ranking assumes everyone in your pool wants the same things. They rarely do, and the averaging hides the disagreements worth knowing about. Splitting by pricing plan, region or seniority is where the useful answer usually lives.

 

Two or three multiple-choice questions at the start make this possible later. After that, tapping any bar on the results page re-ranks everything for that group.

What is the history and origins of pairwise comparison?

The method started in psychology, not in data science or mathematics.

1927.A Law of Comparative Judgment appears in Psychological Review, written by the psychologist L. L. Thurstone. He aims it at physical objects first, ones with properties you can weigh. His larger legacy is elsewhere, in multiple-factor analysis and Primary Mental Abilities, which influenced how intelligence tests came to be structured.

1929. A follow-up, "The Measurement of Psychological Value", proves the same approach scales to intangibles. Attitudes and values become measurable through how strongly people prefer them.

In the same year, the German mathematician Ernst Zermelo applies the idea somewhere else entirely: ranking chess players who haven't all played each other.

1952. Two American academics build on Zermelo and publish the Bradley-Terry model, whose mathematics now underpins sports rankings, peer-review ordering and parts of contemporary machine learning.

1960s. Arpad Elo, a Hungarian-American physics professor, reads Zermelo and builds the chess rating system that carries his name.

1995 and 2005. Glicko, then TrueSkill, extend Elo's approach. Both run today inside Pokémon Go, Chess.com, Dota and Counter-Strike.


Frequently asked questions

What is pairwise comparison in simple terms?

A way of ranking a list by showing people two options at a time and asking which they prefer. Repeat across different pairs and the combined votes produce a full ranking, with a score for each option instead of just a position.

How do you calculate a pairwise comparison matrix?

List your options down the side and across the top, then work along each row comparing that option against every column. Mark 1 for a win and 0 for a loss. Total each row and sort. Ties usually get 0.5 to both sides.

How many pairs are in a pairwise comparison?

n(n-1)/2, where n is the number of options. Ten options produce 45 pairs, 20 produce 190 and 100 produce 4,950. Past about 20 options, switch to a partial comparison where each participant sees a sample.

What's the difference between this and MaxDiff?

Pairwise shows two options per screen, MaxDiff shows three to six and asks for the best and worst. MaxDiff gathers more data per vote, which suits a small participant pool. Pairwise keeps each decision lighter, which suits complex or wordy options.

Who invented pairwise comparison?

L. L. Thurstone published A Law of Comparative Judgment in Psychological Review in 1927, giving psychology a formal model for scaling paired judgments. Ernst Zermelo, the Bradley-Terry model and Arpad Elo's ELO rating all built on that foundation.


Over 42,000 researchers and product people get one method breakdown like this each week in The Full-Stack Researcher.

Setup runs about three minutes. Nothing sits behind the paywall except volume: every question type and every analysis feature works free, up to 25 participants per survey, which is enough to run a real study and read the results before deciding anything.

Previous
Previous

What Is Stack Ranking: Meaning, Examples, Templates, Advice

Next
Next

Web3 Won’t Save Us: How product analytics contaminates decentralized movements