What User Testing Is: Definition, Types and How to Run It
- User testing (usability testing) is watching real people attempt real tasks on your site to see where they struggle; its purpose is to explain the WHY behind behaviour your analytics only hints at.
- It's qualitative, not statistical: it works with small samples (often five to eight people) because it looks for PATTERNS of confusion, not their exact frequency, so a handful of representative users surfaces most of the major issues.
- The main types are moderated (facilitator present, best for the WHY), unmoderated (self-guided, fast and cheap), remote (participants in their own environment), and in-person or guerrilla (deep observation); most teams mix them.
- Run it in five steps: define tasks and goals, recruit a few real users, have them think aloud, watch and note the friction, then turn findings into testable hypotheses.
- The honest limit: user testing finds problems and generates hypotheses but does NOT prove a fix works at scale, for that you run an A/B test. The two are partners, user testing tells you what to fix, A/B testing proves the fix.
Your analytics can tell you that people abandon your checkout, but it will never tell you why. That gap, between knowing what happens and understanding why it happens, is exactly what user testing fills. By watching real people attempt real tasks on your site, you see the hesitation, the misread label, and the moment someone gives up, none of which shows up as a number on a dashboard. This guide explains what user testing is, the main types, how to run a session, how it differs from A/B testing, and how the two work together, drawing on the experimentation practice behind 70,000+ experiments across 7,000+ websites in 15+ industries [CROBenchmark Report 2026, Omniconvert].
The core idea: user testing is a discovery tool that finds problems and generates hypotheses; it doesn't prove a fix works at scale, and that's where A/B testing takes over.
What user testing is
User testing, also called usability testing, is the practice of watching real people attempt real tasks on your site or product to see where they struggle. You give a small number of participants a concrete goal, find a product and add it to the cart, complete checkout, find the return policy, and then observe what they do, where they hesitate, and where they fail. The purpose is not to gather statistics; it's to understand behaviour.
User testing answers the WHY behind what your analytics already shows. Your data might tell you that most people abandon the checkout, but it won't tell you that they abandon because the shipping cost appears only at the final step, or because a required field is confusing. It's qualitative and works with small samples, often five to eight people per round, because the goal is to observe patterns of confusion, not to measure their exact frequency. That single distinction, discovery not measurement, shapes everything about how you choose a format and run a session.
Types of user testing
User testing comes in a few common forms. They differ mainly in whether a facilitator is present and where the session happens, and each fits a different question. The table sets out when each is the right choice.
| Type | How it works | Best for |
|---|---|---|
| Moderated | A facilitator guides the participant live, asking follow-up questions and probing hesitation as it happens | Exploratory research and complex flows where you need to understand reasoning |
| Unmoderated | Participants complete tasks on their own, recorded by a tool, with no facilitator present | Straightforward tasks and larger numbers of sessions, run fast and cheaply |
| Remote | Run over the internet with participants in their own environment, on their own devices | Reaching a geographically spread audience quickly, in realistic conditions |
| In-person (guerrilla) | You sit with the participant in the same room (guerrilla testing approaches people informally in a public place) | Deep observation, capturing body language and richer reactions |
These are not exclusive. Most teams mix them, a couple of moderated sessions to understand the WHY, then unmoderated rounds to confirm the pattern holds, and whichever format you choose, the session itself follows the same steps.
How to run a user test
A user test follows a simple, repeatable process. The point of each step is to keep the session grounded in real behaviour and to end with something you can act on.
- Define the tasks and goals. Decide what you want to learn and write realistic, specific tasks that map to real user goals, such as "buy this product" or "find the delivery timescale", rather than vague instructions.
- Recruit a few real users. Find participants who resemble your actual audience. A handful of representative people, often five to eight, is enough to surface the major issues in one round.
- Have them think aloud. Ask participants to narrate their thoughts as they work, so you hear their expectations, confusion, and reasoning, not just their clicks.
- Watch and note the friction. Observe quietly, avoid leading them, and record every point where they hesitate, backtrack, misread, or fail, noting the patterns that repeat across participants.
- Turn findings into hypotheses. Convert the observed problems into specific, testable statements, such as "moving the shipping cost earlier will reduce checkout abandonment", ready to be validated.
That final step is the hinge. A hypothesis is precisely what an A/B test is built to prove, which is why user testing and A/B testing are best understood as two halves of one process.
User testing vs A/B testing
User testing and A/B testing are often set against each other, but they answer different questions and work best as partners. User testing is qualitative: it finds problems and explains WHY people behave as they do, by observing a small number of real users. It tells you what is broken and why, but because it uses tiny samples, it cannot tell you how common a problem is, or prove that a change actually improves results.
A/B testing is quantitative: it takes a specific change, shows it to a large, randomly split audience, and measures whether it moves a metric to statistical significance. It proves a fix works at scale, but on its own it won't tell you what to change or why an idea might work. So the two run in a natural sequence, user testing tells you what to fix and generates the hypothesis, A/B testing proves the fix. Using only user testing risks shipping changes that feel right but don't move the numbers; using only A/B testing risks testing weak, uninformed ideas. Together they are far more reliable than either alone, which is exactly the pairing an experimentation platform is built to support.
How Omniconvert Explore uses the results of user testing
Omniconvert Explore is an A/B testing and experimentation platform, and it sits at the step directly after user testing: it takes the hypotheses that user testing produces and proves them with controlled experiments at statistical significance. User testing tells you what to fix and why, but on its own it can't prove a fix works for your whole audience, because it observes only a handful of people. Explore closes that gap.
You take a hypothesis from a testing session, "moving the shipping cost earlier will reduce checkout abandonment", and turn it into an A/B test that splits real traffic between the current version and the change, measuring the effect on conversion until the result is statistically trustworthy. Explore also pairs naturally with qualitative tools such as session recordings, so you can watch the friction a test revealed play out across many real visits, and segmentation lets you see whether a change helps specific groups such as new versus returning visitors. Across more than 70,000 experiments, with an average uplift of 23.2%, the reliable pattern is the same: user testing finds the problem, Explore proves the solution.
You've found the friction. Now prove the fix lifts conversion for your whole audience.
See how Omniconvert Explore turns hypotheses into tested wins →Frequently Asked Questions
User testing, also called usability testing, is the practice of watching real people attempt real tasks on your site or product to see where they struggle. You give a small number of participants a concrete goal, find a product and add it to the cart, complete checkout, find the return policy, and then observe what they do, where they hesitate, and where they fail. The purpose is not to gather statistics; it's to understand behaviour. User testing answers the WHY behind what your analytics already shows: your data might tell you that 60% of people abandon the checkout, but it won't tell you that they abandon because the shipping cost appears only at the final step, or because a required field is confusing. User testing is qualitative and works with small samples, often five to eight people per round, because the goal is to observe patterns of confusion, not to measure their exact frequency. That single distinction is the key to using it well: it's a discovery tool that finds problems and generates hypotheses, not a measurement tool that proves a fix works at scale.
There are four common types, and they differ mainly in whether a facilitator is present and where the session happens. Moderated testing has a facilitator guiding the participant live, asking follow-up questions and probing hesitation as it happens; it's best for exploratory research and complex flows where you need to understand reasoning. Unmoderated testing lets participants complete tasks on their own, usually recorded by a tool, with no facilitator present; it's faster and cheaper and best for straightforward tasks and larger numbers of sessions. Remote testing is run over the internet with participants in their own environment, on their own devices; it's best for reaching a geographically spread audience quickly and testing in realistic conditions. In-person testing (sometimes run informally as guerrilla testing, where you approach people in a public place) puts you in the same room as the participant; it's best for deep observation, capturing body language and richer reactions. Most teams mix these: a couple of moderated sessions to understand the WHY, then unmoderated rounds to confirm the pattern holds.
Run a user test in five steps. First, define the tasks and goals: decide what you want to learn and write realistic, specific tasks that map to real user goals, such as buy a specific product or find the delivery timescale, rather than vague instructions. Second, recruit a few real users who resemble your actual audience; a handful of representative people (often five to eight) is enough to surface the major issues in one round. Third, have them think aloud: ask participants to narrate their thoughts as they work, so you hear their expectations, confusion, and reasoning, not just their clicks. Fourth, watch and note the friction: observe quietly, avoid leading them, and record every point where they hesitate, backtrack, misread, or fail, noting patterns that repeat across participants. Fifth, turn findings into hypotheses: convert the observed problems into specific, testable statements such as moving the shipping cost earlier will reduce checkout abandonment, ready to be validated. That last step is what connects user testing to the rest of your optimisation work, because a hypothesis is exactly what an A/B test is built to prove.
User testing and A/B testing answer different questions and are best used together. User testing is qualitative: it finds problems and explains WHY people behave as they do, by observing a small number of real users attempt tasks. It tells you what is broken and why, but because it uses tiny samples it cannot tell you how common a problem is or prove that a change actually improves results. A/B testing is quantitative: it takes a specific change, shows it to a large, randomly split audience, and measures whether it moves a metric to statistical significance. It proves that a fix works at scale, but on its own it won't tell you what to change or why an idea might work. So the two are partners in a natural sequence: user testing tells you what to fix and generates the hypothesis; A/B testing proves the fix. Using only user testing risks shipping changes that feel right but don't move the numbers; using only A/B testing risks testing weak, uninformed ideas. Together, user testing points you at the right problems and A/B testing confirms the solutions, which is a far more reliable way to improve a site than either alone.
For qualitative user testing you need far fewer people than most teams expect, typically around five to eight participants per round. The reason is that user testing looks for patterns of confusion, not their exact frequency, and a small group of representative users tends to surface the large majority of the significant usability problems in a given flow. Once several people hit the same obstacle, you have learned what you needed to learn about that obstacle; adding many more participants mostly reveals the same issues again, with diminishing return. This is why user testing is best run in small, frequent rounds rather than one large study: test a handful of people, fix what you found, then test again. It's important not to confuse this with the sample sizes needed for A/B testing, which are much larger. A/B testing is statistical, so it needs enough visitors per variation to reach confidence, often thousands. User testing is observational, so a few well-chosen participants are enough. The two use small and large samples for completely different, and complementary, reasons.
User testing is the discovery stage of conversion rate optimisation (CRO). A disciplined CRO process starts by finding out where and why people struggle, and user testing is one of the strongest ways to gather that WHY, alongside session recordings, heat maps, surveys, and heuristic (expert) evaluation. These qualitative methods reveal the friction points that raw analytics numbers only hint at, and they feed the hypothesis backlog: each observed problem becomes a specific, testable idea for improvement. From there, the quantitative stage takes over: you prioritise the hypotheses, then run A/B tests to prove which changes genuinely lift conversion for your whole audience. Skipping the discovery stage means testing guesses, which wastes traffic on ideas that were never grounded in real user behaviour. So user testing doesn't replace experimentation; it makes experimentation smarter, by ensuring the changes you test are aimed at real problems real people are having. In a mature CRO programme, user testing and A/B testing run continuously, feeding each other: observe, hypothesise, test, learn, repeat.
Omniconvert Explore is an A/B testing and experimentation platform, and it sits at the step directly after user testing: it takes the hypotheses that user testing produces and proves them with controlled experiments at statistical significance. User testing tells you what to fix and why, but on its own it can't prove a fix works for your whole audience, because it observes only a handful of people. Explore closes that gap. You take a hypothesis from a testing session, moving the shipping cost earlier will reduce checkout abandonment, and turn it into an A/B test that splits real traffic between the current version and the change, measuring the effect on conversion until the result is statistically trustworthy. Explore also pairs naturally with qualitative tools such as session recordings, so you can watch the friction a test revealed play out across many real visits before and after you fix it, and segmentation lets you see whether a change helps specific groups such as new versus returning visitors or mobile users. Across more than 70,000 experiments, with an average uplift of 23.2%, the reliable pattern is the same: user testing finds the problem, Explore proves the solution.
User testing is watching real people attempt real tasks on your site to see where they struggle. It's qualitative, works with small samples of five to eight, and answers the question analytics can't: WHY people behave as they do. That makes it a discovery tool, it finds problems and generates hypotheses, not a measurement tool. This is the honest limit worth remembering: user testing tells you what is broken and why, but it does not prove that a fix works at scale, because a handful of people can't tell you how a change lands with your whole audience. For that you need an A/B test. The two are partners, not rivals: user testing tells you what to fix, A/B testing proves the fix. Run them in sequence, observe and hypothesise with user testing, then validate with experimentation, and you avoid the two classic traps: shipping changes that feel right but don't move the numbers, and testing weak ideas that were never grounded in real behaviour. That partnership is exactly how Omniconvert Explore is built to be used: it takes the hypotheses user testing produces and proves them with controlled experiments at statistical significance.
Prove the fixes user testing finds, with Omniconvert Explore
Omniconvert Explore takes the hypotheses your user testing produces and proves them with A/B tests at statistical significance, so you ship changes that genuinely lift conversion, not just changes that feel right. Pair it with session recordings to watch the friction play out, and segment results by new vs returning, device, or source.