Surveys

How to Choose the Right Survey Scale (12 Types Compared)

First published Feb 20, 2023Updated September 7, 202616 min read
Valentin Radu, Founder and CEO of Omniconvert
Valentin Radu
Founder & CEO, Omniconvert · Author, The CLV Revolution
Published: Feb 20, 2023Updated: Sep 7, 2026
Reviewed by Cristina Stefanova, Head of Content
Customer feedback kiosk with sad, neutral, and blue happy face buttons
Quick Answer
A survey scale is the ordered set of response options a respondent chooses from, and the right one is the one that matches the answer you need. Use a Likert scale to measure agreement with a statement, a semantic differential scale to place something between two opposite adjectives, a numeric or star scale for a quick rating, NPS (always 0 to 10) for likelihood to recommend, and a binary scale only when the answer really is yes or no. Five to seven points suits most attitude questions: an odd number of points offers a neutral midpoint, an even number forces respondents to lean one way. Decide what you will do with the answer before you decide the scale, because a scale that cannot produce a decision is a wasted question. Surveys in Omniconvert Explore run these scales on-site and post-purchase and feed the answers into the experiments you run next.
Key Takeaways
  • The scale defines what the answer can say. Pick the scale from the decision you need to make, not from what looks tidy on the page.
  • An odd number of points gives a neutral midpoint; an even number forces a lean. Neither is more valid, but a midpoint also collects the indifferent, the uninformed and the unwilling, and you cannot tell them apart afterwards.
  • Five to seven points fits most attitude questions. Below five loses detail, above seven asks for distinctions people cannot reliably make.
  • NPS is a 0-to-10 scale, not 1 to 10. Fred Reichheld defined it that way in Harvard Business Review in 2003, and changing the range breaks the promoter, passive and detractor bands.
  • Label every point on a Likert scale, but label only the endpoints on a semantic differential. Mixed or uneven labelling biases the answer before the respondent reads the question.
7,000+ websites using Omniconvert 70,000+ experiments run 23.2% average conversion uplift 15+ industries covered

A survey scale is the ordered set of response options a respondent chooses from, and the right one is the one that matches the answer you need. Use a Likert scale for agreement with a statement, a semantic differential scale to place something between two opposite adjectives, a numeric or star scale for a quick rating, NPS (always 0 to 10) for likelihood to recommend, and a binary scale only when the answer genuinely is yes or no.

Nothing is more frustrating than caring about a subject and being offered only "yes" and "no". The reverse is just as bad: having a simple answer and being made to write thirty words. Both are scale problems, and both are decided before a single response arrives. This article covers the twelve scales worth knowing, how many points to give them, how to label them, and how to pick one.

What a survey scale is, and why the choice matters

A survey scale is the ordered set of response options a respondent picks from to express how much of something they feel: agreement, satisfaction, frequency, importance or likelihood. The scale turns an opinion into a value you can count, compare across respondents and track over time. Choosing a scale is choosing what the answer can say, so the scale has to match the question.

The scale does two jobs at once. It makes answers comparable, so a thousand opinions become a distribution you can read, chart and re-measure next quarter, whether in Excel or in a statistics package such as SPSS. And it tells the respondent what kind of answer you are asking for. A five-point agreement scale signals "give me a considered position". A single star tap signals "one second is all I need".

The second job is the one people underestimate. Words that are obvious to the person writing the survey are not always obvious to the person answering it. Neuro-Linguistic Programming (NLP) describes this as every person having their own "mind map". "Satisfied" means one thing to a researcher and another to a customer whose parcel arrived two days late but intact. Most inaccurate surveys are not badly analyzed; they are answered by people who understood the question differently from the person who wrote it. A well-chosen, well-labelled scale narrows that gap. It is the cheapest accuracy you will ever buy.

It also constrains what you can do afterwards. A binary answer can be counted but not averaged. A ranking tells you order but not distance. A five-point Likert scale is ordinal, so the honest summary is the distribution and the median, not a mean to two decimal places. Decide what the analysis has to look like before the scale is set, because no amount of work afterwards recovers detail the scale never captured.

The 12 survey scale types, compared

The twelve survey scales in common use are Likert, semantic differential, linear numeric, NPS, star, binary, pictorial, matrix, forced ranking, paired comparison, constant sum and adjective checklist. The first seven measure the strength of an opinion on one item. The last five compare items against each other and answer questions about priority rather than intensity.
Survey scale types and what each one is for
Scale What it measures Typical format Use it when
Likert Strength of agreement with one statement 5 or 7 points, every point labelled You want a considered position on a specific claim
Semantic differential Position between two opposite adjectives 5 to 7 points, endpoints labelled only You are measuring perception of a brand, product or price
Linear numeric Level or degree of one attribute 1-5, 1-7 or 1-10, endpoints anchored You want a fast rating you can track over time
NPS Likelihood to recommend 0-10, fixed; 9-10 promoters, 7-8 passives, 0-6 detractors You need a comparable loyalty benchmark after a real experience
Star Overall satisfaction, in a familiar shorthand 1-5 stars Attention is short and the rating will be shown publicly
Binary Presence or absence of one thing Yes/no, did/did not The answer really is two-valued, or you need a routing question
Pictorial / graphic Emotional reaction, without reading Faces or icons, usually 3 to 5 Respondents are passing by, or may not read the survey language well
Matrix The same scale applied to several items at once Grid of items x one shared scale You are rating a set of related attributes and want them side by side
Forced ranking Order of preference, no ties allowed Drag or number up to about 10 items You need to know what comes first, not how much each is liked
Paired comparison Preference between exactly two options Repeated A vs B choices You have a short list already and need a clean winner
Constant sum Relative weight of each item Allocate 100 points or 100 dollars across items You need the size of the gaps between priorities, not just the order
Adjective checklist Which associations people hold Pick all that apply from a mixed positive and negative list You are mapping brand perception rather than measuring one attitude

Rating scales: measuring one thing at a time

Rating scales measure how much of one attribute a respondent perceives. Likert measures agreement with a statement, semantic differential places a concept between two opposite adjectives, linear numeric and star scales capture a quick level, NPS captures likelihood to recommend on a fixed 0-to-10 range, binary captures a two-valued fact, and pictorial scales capture a reaction without asking anyone to read.

Likert scale

The Likert scale asks the respondent to agree or disagree with a single statement, one statement at a time. Rensis Likert introduced it in 1932 in A Technique for the Measurement of Attitudes, using five points from strongly approve to strongly disapprove. Five points is still the classic form; seven is common when you want more resolution, and every point carries a written label.

Because it says "do you agree", it reads as personal, and respondents tend to answer it more carefully than a bare number. Use it when you want a considered position on a specific claim: "The delivery options were clear before I paid." Write one idea per statement, or you will not know which half of it the answer refers to. Formplus collects Likert scale examples if you want to see the variants.

Semantic differential scale

A semantic differential scale puts two opposite adjectives at the ends of a line and asks the respondent to mark a point between them: cheap to expensive, complicated to simple, cold to friendly. Charles Osgood, George Suci and Percy Tannenbaum introduced the technique in The Measurement of Meaning in 1957 to measure the connotations people attach to a concept. Only the endpoints are labelled.

Use it for perception rather than agreement. It is the right instrument when you want to know how a brand, a product page or a price feels, and when the interesting answer is a position between two poles rather than a level of one attribute.

Linear numeric scale

A linear numeric scale offers a level of something — satisfaction, effort, interest, clarity — usually 1 to 5, 1 to 7 or 1 to 10, with the endpoints anchored in words ("not at all" to "extremely"). It is a type of semantic differential scale (QuestionPro explains the format in detail). It is the workhorse of quick feedback and the basis of the standard customer experience metrics. CSAT and CES are both linear numeric scales with fixed wording.

Use it when you want a number you can track wave over wave and you do not need to know why. Keep the direction consistent across every numeric question in the survey; flipping high and low mid-survey is one of the most common sources of dirty data.

NPS (0 to 10)

The Net Promoter Score scale is a fixed 0-to-10 range, eleven points, asking how likely the respondent is to recommend you. Fred Reichheld introduced it in the 2003 Harvard Business Review article The One Number You Need to Grow, phrased "on a zero-to-ten scale". Scores of 9 and 10 are promoters, 7 and 8 are passives, and 0 through 6 are detractors.

The range is not a design choice. Running NPS as 1 to 10, or as 0 to 5, shifts the bands and makes your score incomparable with everyone else's, which removes the only real reason to use NPS instead of a plain satisfaction rating. If you want to change the scale, you want a different metric.

Star rating

A star rating is a linear numeric scale wearing a familiar costume, almost always 1 to 5. Its advantage is that no one needs the instructions. Its limitation is the same one that applies to public review scores: ratings cluster at the top and at 1, so the average moves slowly and hides a lot. Use it where the rating will also be displayed, or where a single tap is genuinely all you will get.

Binary scale

Yes or no, did or did not. A binary scale is the correct instrument when the underlying fact is two-valued ("Did you find what you were looking for?") and the wrong one whenever the honest answer is "partly". It is also useful as a routing question that decides which scale a respondent sees next.

Pictorial and graphic scales

Faces, icons or colored expressions replace the words: a green smile at one end, a red frown at the other. These are the scales in airport corridors and hospital corridors, and they work for the same reason in both places. People answer them in passing, without reading, and they survive language barriers. Use them where response rate matters more than nuance, and accept that three to five levels is all they can carry.

Run NPS, CSAT, rating and open-ended surveys on-site and after purchase, targeted by segment.

See Surveys in Omniconvert Explore →

Comparative and preference scales: measuring what comes first

Comparative scales make respondents weigh items against each other instead of rating each one on its own. Matrix scales apply one shared scale to several items, forced ranking puts items in order with no ties, paired comparison reduces the choice to two options at a time, constant sum asks respondents to split 100 points across items, and an adjective checklist collects associations rather than levels. Use them when the question is about priority, not intensity.

Rating scales have a well-known failure mode: ask people to rate ten features and most of them rate eight as important. That is not a lie, it is what happens when nothing forces a trade-off. Comparative scales force one.

  • Matrix scale. A grid of items sharing one scale, so several related attributes are rated side by side and stay comparable. Efficient and compact, but it turns tedious fast, so keep the row count low and never stack two matrices in a row.
  • Forced ranking. Respondents order up to about ten items with no ties permitted. Stronger than individual ratings precisely because it removes the option of calling everything important. Randomize the starting order for each respondent, or the items shown first will drift to the top.
  • Paired comparison. Two options at a time, repeated. Easier for the respondent than ranking a long list, and it produces a clean winner. Use it after a ranking round has already narrowed the field; used first, it can lock you into two options that were never the important ones.
  • Constant sum. Respondents split 100 points, or 100 hypothetical dollars, across the items. It captures the size of the gaps that ranking flattens: first place with 60 points is a very different finding from first place with 22. Keep the list short, because the arithmetic is real work.
  • Adjective checklist. A mixed list of positive and negative adjectives, from which respondents pick everything they associate with the brand or product. Not a scale in the strict sense, but it belongs here because the output is comparative. Include adjectives you would not choose yourself, randomize the order, and keep the positive and negative counts balanced.

Odd or even? How many points a scale should have

Use an odd number of points when a genuinely neutral position is a real answer, because the midpoint gives respondents somewhere honest to sit. Use an even number, a forced-choice scale, when you want every respondent to lean one way. Five to seven points fits most attitude questions: below five loses detail, above seven asks for distinctions people cannot reliably make.

The odd-versus-even decision is about the midpoint, and it is a real trade-off rather than a best practice.

  • An odd scale (3, 5, 7) has a neutral midpoint. That is honest for respondents who truly have no view, and it prevents a fabricated lean. The cost is that the midpoint becomes a catch-all: the indifferent, the uninformed and the unwilling all land on it, and afterwards you cannot tell those three groups apart.
  • An even scale (4, 6) forces a side. Every respondent commits, and the results split cleanly. The cost is that people with a genuine neutral position are pushed off it, so a small artificial lean is baked into the data. Choose it when a decision needs a direction, not when you are measuring a true attitude.

If you go odd, label the midpoint honestly — "neither agree nor disagree" — and give a separate "don't know" or "not applicable" option outside the scale. That single addition is what stops the midpoint from becoming a shrug. If your survey tool cannot offer it, prefer an even scale over a midpoint that means four different things.

On length: five to seven points covers almost every attitude question. Three points is too coarse to show movement between waves. Beyond seven, respondents cannot reliably distinguish an 8 from a 9, so the extra points add noise rather than resolution and different people use different parts of the range. Fixed-format metrics are the exception and keep their defined length whatever you think of it: NPS stays at 0 to 10, star ratings stay at 1 to 5.

Labelling and wording the scale

Label every point on a Likert scale, because the labels are what the respondent is answering. Label only the endpoints on a semantic differential or a long numeric scale, where the numbers already carry the order. Keep the wording evenly spaced, keep the same number of positive and negative options, and never change direction partway through a survey.

Labels are not decoration. On a fully labelled scale the respondent is choosing a phrase, not a number, so uneven phrases skew the result. "Excellent, very good, good, fair, poor" looks balanced but is not: four of the five options sit at or above neutral. A balanced set has as many negative steps as positive ones, with steps that feel evenly spaced.

  1. Match the label set to the question stem
    Agreement questions get agree/disagree labels. Satisfaction questions get satisfaction labels. Frequency questions get frequencies with real intervals ("weekly", not "often"). Mixing stems and labels is the most common reason answers look inconsistent.
  2. Balance positive and negative options
    Equal counts on each side of the midpoint. If the negative half has fewer or milder options, the average moves up on its own and the survey flatters you.
  3. Keep one direction for the whole survey
    If low is bad in question one, low is bad in question ten. Flipping polarity to "keep respondents attentive" mostly produces mistakes you cannot detect afterwards.
  4. Anchor the endpoints in plain words
    On a numeric scale, "1 = not at all likely, 10 = extremely likely" is doing the real work. A bare row of numbers means different things to different people.
  5. Put "don't know" outside the scale
    A no-opinion option belongs next to the scale, not inside it. Inside, it corrupts the midpoint and any average you calculate from it.

How to choose a survey scale, step by step

Choose a survey scale by working backwards from the decision the answer must support. Name the decision, name the construct you are measuring, pick the scale family that produces that kind of answer, set the number of points and the labels, then check that the scale can be repeated unchanged for as long as you want to compare results.
  1. Name the decision the answer has to support
    "We will change the shipping copy if people find it unclear" is a decision. "We want to understand our customers" is not. A question that cannot change anything does not need a scale; it needs deleting.
  2. Name the construct you are measuring
    Agreement, satisfaction, effort, likelihood, frequency, importance or preference. Each one has a scale family that fits it. Getting this wrong is why so many surveys measure agreement when they meant to measure priority.
  3. Pick rating or comparative
    If the answer is about how much, use a rating scale. If it is about which comes first, use ranking, paired comparison or constant sum. Rating scales cannot settle a prioritization question, no matter how the results are analyzed afterwards.
  4. Set the number of points
    Five to seven for attitudes. Odd if neutral is a real answer, even if you need a lean. Keep the fixed formats fixed: NPS 0 to 10, stars 1 to 5.
  5. Write the labels and read them aloud
    Balanced, evenly spaced, one direction, endpoints anchored. If a label sounds awkward when spoken, it will be misread on screen.
  6. Pilot it on a small sample before the full launch
    Twenty to fifty responses will show you the failure modes: everyone picking the midpoint, no one using the bottom half, a label people clearly read two ways. Fix them before the scale is locked in.
  7. Freeze it for as long as you want to compare
    Once responses start arriving, the scale is part of the measurement. If you must change it, close the wave and report the two periods separately rather than pooling them.

Scale mistakes that bias the answers

The most common survey scale mistakes are double-barrelled questions, leading wording, unbalanced label sets, too many points, a midpoint that doubles as a no-opinion option, and asking a scale question at a moment when the respondent has no basis for answering. Each one produces data that looks clean and is not.
  • Two questions in one. "The site was fast and easy to use" cannot be answered by anyone for whom only half is true. One idea per statement.
  • Leading wording. "How much did you enjoy the new checkout?" has already decided that you enjoyed it. Ask neutrally and let the scale carry the direction.
  • Metaphors, jargon and puns. Basic vocabulary answers faster and more accurately. A clever question is a slower question.
  • Unbalanced labels. More positive options than negative ones raises the score without changing anything real.
  • Too many points. An eleven-point agreement scale does not produce eleven meaningful levels. It produces five levels and some noise.
  • A survey longer than it needs to be. Response quality falls as length rises, and the answers at the end are worse than the answers at the start. Cut every question that does not feed a decision.
  • Asking at the wrong moment. An NPS question shown before delivery measures expectation, not experience. Timing is part of the scale design, not a separate concern.

If you want a starting point rather than a blank page, Jotform keeps a large gallery of survey templates.

Which scale to use where in the customer journey

Match the scale to the attention available at that moment. On-site, use one short numeric, star or binary question. Immediately after purchase, use a short rating on the buying experience. After delivery, use NPS or CSAT, once the customer has something real to judge. In a recruited research survey, use Likert, matrix or comparative scales, where the respondent has already agreed to spend a few minutes.

The same question gets a different answer depending on where it is asked. Feedback at each key point of the journey is worth more than a single long survey at the end, because the respondent still remembers the specific moment you are asking about. Our guide to the four survey types for eCommerce covers the placement side in detail, and the survey questions guide covers the wording.

Source: Omniconvert
Moment Attention available Scale that fits What the answer is good for
On a product page Seconds, mid-task Binary or 1-5 numeric, one question Finding the specific page objection to test next
Exit intent, before leaving Seconds, low goodwill Binary plus one optional open field Naming the reason people leave without buying
Immediately after checkout Short, goodwill high CES or a 1-5 rating of the buying experience Locating friction in the purchase flow while it is fresh
After delivery Moderate, experience complete NPS (0-10) or CSAT Tracking loyalty and satisfaction wave over wave
Recruited research survey Minutes, agreed in advance Likert, matrix, semantic differential Understanding attitudes and perception in depth
Roadmap or pricing research Minutes, engaged respondents Forced ranking, paired comparison, constant sum Settling priority questions a rating scale cannot settle

Surveys in Omniconvert Explore run these formats on-site and post-purchase, trigger them by behavior and segment, and connect the answers to the experiments that follow. That last part is the point: a scale earns its place when the answer changes something. Explore turns a survey finding into an A/B test against the same audience, so the opinion you collected gets checked against behavior.

Ask your visitors and customers directly, then test what they tell you. FREE A/B testing on 50,000 visitors with Omniconvert Explore.

Start for free →

Frequently Asked Questions

1What is a survey scale?

A survey scale is the ordered set of response options a respondent picks from to express how much of something they feel: agreement, satisfaction, frequency, importance or likelihood. The scale turns an opinion into a value you can count, compare across respondents and track over time. Choosing a scale is choosing what the answer can say, so the scale has to match the question.

2Should a survey scale have an odd or an even number of points?

Use an odd number of points when a genuinely neutral position is a real answer, because the midpoint gives respondents somewhere honest to sit. Use an even number, a forced-choice scale, when you want every respondent to lean one way and you accept that some will be pushed off a true neutral. Neither is more correct. The risk with an odd scale is that the midpoint collects people who are indifferent, uninformed or unwilling to answer, and you cannot tell those three apart afterwards.

3How many points should a rating scale have?

Five to seven points fits most attitude questions. Fewer than five throws away detail and makes small changes invisible between waves. More than seven asks people to make distinctions they cannot reliably make, so the extra points add noise rather than precision. Fixed-format metrics are exceptions and keep their own defined length: NPS is always 0 to 10, and star ratings are almost always 1 to 5.

4What is the difference between a Likert scale and a semantic differential scale?

A Likert scale asks how strongly a respondent agrees or disagrees with a single statement, so every point carries a written label from strongly disagree to strongly agree. A semantic differential scale puts two opposite adjectives at the ends of the line, such as cheap and expensive, and only the endpoints are labelled. Likert measures the strength of a position on a statement. Semantic differential measures where a brand or product sits between two poles.

5Is the NPS scale 0 to 10 or 1 to 10?

The Net Promoter Score scale is 0 to 10, giving 11 points. Fred Reichheld introduced it in the 2003 Harvard Business Review article The One Number You Need to Grow with the wording "on a zero-to-ten scale". Scores of 9 and 10 are promoters, 7 and 8 are passives, and 0 through 6 are detractors. Running it as 1 to 10 breaks the promoter, passive and detractor bands and makes the result impossible to compare with anyone else's NPS.

6Do you have to label every point on a survey scale?

Label every point on a Likert scale, because the labels are what the respondent is answering. Label only the endpoints on a semantic differential or a long numeric scale, where the labels would crowd the line and the numbers already carry the order. Whichever you choose, keep the spacing of the wording even, keep the same number of positive and negative options, and never mix labelled and unlabelled points in the same question.

7Can I change a survey scale after the survey has started?

You can, but the responses collected before and after the change are not comparable and should not be pooled or averaged together. Changing the number of points, the wording of the labels, or the direction of the scale changes what the numbers mean. If a scale has to change, close the current wave, start a new one, and report the two separately with the change noted.

8Which survey scale should I use on an eCommerce site?

Use a short numeric or star scale on-site, where attention is scarce and a single tap is all you will get. Use NPS after delivery, once the customer has actually experienced the product. Use Likert or matrix scales in a longer post-purchase or research survey where the respondent has already agreed to spend a few minutes. Use ranking or constant-sum scales only for research questions about priorities, never as a passing on-site question.

The short version

Write down the decision the answer has to support before you choose a scale. If the decision is "which of these five features do we build first", a rating scale will not settle it and a ranking or constant-sum scale will. If the decision is "did the checkout change make people feel better about paying", a labelled five-point agreement scale will. Then keep the scale fixed for as long as you want to compare waves, label it evenly, and put it where the respondent has just had the experience you are asking about. A survey is only as accurate as the scale allows it to be, and no amount of analysis afterwards can recover detail the scale never captured.

Valentin Radu, Founder and CEO of Omniconvert
Founder & CEO, Omniconvert
Valentin Radu is the founder and CEO of Omniconvert. He is an entrepreneur, data-driven marketer, CRO expert, CVO evangelist, international speaker, father, husband, and pet guardian. Valentin is also an Instructor at the Customer Value Optimization (CVO) Academy, an educational project that aims to help companies understand and improve Customer Lifetime Value.

Ask better questions on-site and after purchase

Surveys in Omniconvert Explore run NPS, rating, Likert and open questions on-site and post-purchase, target them by segment and behavior, and feed the answers straight into the A/B tests you run next. Explore has run 70,000+ experiments across 7,000+ websites and 15+ industries, with an average conversion uplift of 23.2%.