Website usability testing means watching people attempt realistic tasks on a website and examining where the interface helps or obstructs them. A useful session produces more than a list of opinions. It gives you a traceable account of the task, the participant’s actions, the outcome, any help you supplied and the questions that remain.
For a Canadian business, that might mean observing whether someone can check a service area, understand what a quotation includes or recover from a form error. The immediate goal is a better decision about a specific interaction. A few participants cannot establish how common a problem is across your audience, prove that you found every serious issue or predict a conversion increase.
This guide provides an original working protocol, a task script and an observation rubric. The fictional business, participants, findings and proposed fixes are teaching examples. No participant study, screen recording, customer result or conversion experiment is being reported.
Start with the decision the study must support
“Find out whether people like our website” is too broad to guide recruitment or development. A more useful question is: “Can a first-time visitor determine whether we serve their building and prepare the right information before requesting a quotation?” That question defines who to invite, where the task begins and what observable evidence would support a repair.
Write a one-paragraph study brief before drafting the script. Name the decision owner, the journey under examination, the version of the website and the decision you expect to make. Also name the questions this round cannot answer. A study of an enquiry journey does not establish whether the service is competitively priced or whether your advertising reaches the right market.
The UK Government Digital Service recommends agreeing actionable research objectives and planning analysis alongside the sessions. Those planning principles help keep a small round focused. The specific templates and decision rules below are proposed working conventions for this guide, rather than a validated research instrument. See the GDS guidance on planning a research round.
Imagine a fictional business called Harbour Window Care. It wants to improve an existing quotation journey, not rebuild its whole website. Its decision is whether to change service-area information, clarify the request form and repair the confirmation step. A content lead owns the information, a developer owns the form behaviour and the operations lead confirms what a request actually triggers.
If the work sits inside a wider rebuild, use the website redesign checklist for migration and launch planning. Keep this study’s brief about observed task performance. A list of redesign requirements and a record of a person trying to use the site serve different decisions.
Choose observation for a question about use
A moderated usability session is suitable when you need to see how someone approaches a task and clarify an unexpected action. You present a situation, allow an attempt and ask neutral follow-up questions. The website can be a working site, a test environment or a prototype, provided you describe what it can actually do.
Buying interviews explore experiences such as why someone sought a service, how they compared suppliers and what affected a purchase decision. Those are valuable questions, but they do not replace observing an attempt. A person saying they would use a quotation form is not evidence that they can locate it, interpret its fields or recognise a successful submission.
An expert review is different again. A reviewer may identify a confusing label or a likely keyboard barrier without recruiting anyone. Label that as a review finding. Do not convert an expert’s prediction into a participant observation. Similarly, analytics can indicate where a recorded journey ends but usually need further investigation to explain the cause.
Use usability work to develop and inspect possible changes within a broader measurement plan. Canada Create’s introduction to conversion rate optimisation explains that wider commercial context. The observation protocol here does not calculate lift, estimate revenue or replace a properly designed quantitative evaluation.
For the first Harbour round, moderated sessions with notes are a reasonable planning choice because the team wants to inspect misunderstandings. No recording is necessary for the fictional workflow. In a real project, choose the least intrusive evidence collection that still lets the team make the stated decision.
Recruit for task fit and describe the gaps
Recruitment should follow the journey. If the website is intended for people arranging services for small commercial buildings, inviting only friends who manage household purchases leaves an important mismatch. A convenient participant can help you practise the script; that does not automatically make their attempt evidence about the intended audience.
Create a short recruitment brief using behaviours and circumstances relevant to the task. For Harbour, potential fit includes having helped arrange building maintenance, using a phone for supplier research and having enough responsibility to assess whether a quotation request is appropriate. Do not require knowledge of Harbour’s menu labels or the correct answer to a task.
Useful screening questions could include “What part have you played in arranging building maintenance?” and “Which device would you normally use to check a new supplier?” Ask only for information needed to establish fit or provide access. The aim is not to collect a detailed biography. Keep recruitment contact details outside the observation sheet.
Record whether participants are new to the site, existing customers or already familiar with the organisation. Familiarity can change how they navigate. Staff may know where information lives because they helped create it. If you include them for a particular reason, explain that reason and report their results separately from unfamiliar visitors.
For a bilingual journey, identify the language being examined and prepare tasks that make sense in that language. Completing an English task does not establish that the French journey works. Translation quality, terminology and local service information may need their own review and participant coverage. Do not infer someone’s language needs from their name.
Ask participants what would make participation workable: timing, breaks, communication format, assistance or a familiar device. Keep compensation arrangements clear and independent of positive feedback. Where recruitment depends on a manager or another authority figure, consider whether the person can decline comfortably. A polite agreement under pressure is poor groundwork for candid observation.
Finish the recruitment brief with a coverage statement: who is included, who is absent and why those absences matter. If the round covers only returning desktop users, the report must not quietly become a judgement about first-time mobile visitors. A narrow but explicit study is more useful than an expansive claim supported by mismatched participants.
Plan a small round without making a five-person promise
There is no universal participant count that guarantees useful coverage of every website. Choose a practical initial round around the decision, the variety of journeys, participant differences and the consequences of missing a problem. Reserve time to repair issues and run another round. Treat the initial count as a planning choice, not a certificate of completeness.
Original research helps explain the caution. In Laura Faulkner’s 2003 study, researchers tested 60 users and examined randomly selected subsets. The reported problem discovery varied substantially between different subsets of five. The public abstract illustrates why one small group cannot guarantee a fixed share of problems found; its findings are not a coverage forecast for your website. Read Faulkner’s original research abstract.
James R. Lewis’s 1994 study also qualified small-sample claims: discovery depended on how likely problems were to be encountered, and the data did not support assuming that severe problems would necessarily appear first. That is a reason to consider difficult or consequential paths deliberately. It is not a reason to apply a borrowed discovery percentage to your own study. See Lewis’s original study abstract.
For Harbour, a four-participant example keeps the teaching dataset readable. Four is not a recommendation for all projects. A real team might need additional participants or distinct rounds for different building types, languages, devices or access needs. If one participant represents an important situation, their attempt can reveal a problem but cannot represent everyone in that situation.
Decide in advance what will trigger more investigation: a serious unresolved barrier, conflicting observations, an important untested path or a recruitment gap. Repeating the same easy tasks with similar people may add little understanding of a different journey. Record what the next round needs to learn, rather than declaring the whole site “validated”.
Plan consent, data handling and safe task boundaries
Before inviting anyone, decide what information the study needs, who will see it, where it will be kept and when it will be removed. Notes can contain personal information even when there is no recording. An organisation name, an unusual job description or a detailed account of a previous purchase may identify someone indirectly.
Canadian privacy regulators’ meaningful-consent guidance emphasises understandable explanations of collection, purposes, sharing and meaningful risks. It also notes that organisations must understand the obligations that apply to them. Use that guidance to inform your process; a short research script does not establish compliance with every applicable law. Read the Office of the Privacy Commissioner of Canada’s consent guidance.
For a notes-only session, prepare a plain-language explanation naming the organiser, the study purpose, who will attend, what notes will contain, the intended use, access arrangements and retention dates. Explain how to stop, skip a task or request withdrawal, including any practical limits once information has been irreversibly de-identified. Provide a real contact route and time for questions before seeking agreement. State where notes will be stored, how access will be restricted, any meaningful privacy risks and safeguards, and what happens to compensation if someone stops. Participant codes are pseudonyms, not proof of anonymity; a code key or distinctive details may still identify someone. Explain withdrawal from the session separately from removing data already collected, and do not promise removal you cannot carry out.
Do not silently add a recording or automated transcription service because it is convenient. If the collection method changes, revisit the information and permission process. The same applies to adding observers, sharing extracts beyond the agreed team or using study material in public marketing. Participation in research does not supply permission for a testimonial.
Harbour’s example uses participant codes such as F01 and entirely fictional building details. Its prototype cannot send an enquiry, charge a card or reserve a real appointment. The study plan explicitly excludes actual submissions. In a real preparation check, verify that boundary: a staging page can still send email if someone connected it to production services.
If a participant begins entering personal credentials or confidential customer information, interrupt promptly. A neutral observation protocol does not require watching preventable disclosure. Record that the task was stopped for the planned boundary, not that the participant “failed”. Keep identifiable scheduling and consent records out of any public companion file or shared issue ticket.
Make participation accessible and keep conformance separate
People should be able to participate using an appropriate communication format and interaction method. Ask about access requirements without making someone disclose unnecessary medical details. Test the invitation, joining instructions and research materials as part of preparation; a usable website session can still be inaccessible if the meeting or task sheet creates a barrier.
W3C’s Web Accessibility Initiative recommends involving people with disabilities and cautions against generalising one person’s experience to others. Its guidance also explains that observing users alone cannot determine whether a website is accessible. Combine participant involvement with appropriate standards-based evaluation. See WAI’s guidance on involving users in evaluation.
For example, if a person encounters trouble while using a screen reader, preserve the interaction details and relevant browser and assistive-technology context. Do not assume the cause before investigating. The page, browser, assistive technology, task instructions or their interaction may need examination. Respect the participant’s normal working method rather than demanding a demonstration arranged for observers’ convenience.
Usability observation and accessibility-conformance testing should have separate outcomes. The first describes what happened in particular attempts. The second evaluates specified accessibility requirements within a defined scope and needs relevant expertise. W3C describes WCAG-EM as an approach to conformance evaluation. A successful task attempt is not a substitute for that evaluation. Read WAI’s conformance evaluation overview.
In the report, state the access contexts actually included and leave unexamined contexts visible. Do not write “accessible to everyone” after a small round. A suspected access barrier can still deserve urgent action even if one person encountered it; frequency in the study is not the same as seriousness.
Write tasks that reveal choices without supplying the answer
A task needs a believable situation and a goal. It should not tell the participant which menu item to choose, which phrase to search for or which field contains the answer. Asking someone to “click Service Areas and find the Toronto coverage statement” tests whether they can follow your instruction. It conceals whether they would find the information independently.
GDS’s moderated-testing guidance recommends relevant tasks with clear goals that avoid giving away the route. The original task cards below apply that principle to the fictional Harbour journey. Keep the participant-facing prompt separate from the facilitator’s success criteria and rescue instructions. See the moderated usability testing guidance.
Task T1: assess service eligibility
Read to the participant: “You are arranging window cleaning for a three-storey office building in Toronto. Use this website to decide whether this business is a suitable company to contact. Show what information supports your decision.”
Facilitator criterion: The participant finds and accurately interprets the prototype’s location and building-height limits. They can explain whether the fictional building fits. Reaching the service page alone is insufficient. A correct “not suitable” conclusion would count as completion if the displayed policy excluded that building.
Task T2: prepare a quotation request
Read to the participant: “You have decided to ask for a quotation for this fictional building. Using the supplied sample details, prepare the request as you normally would. Stop before sending anything.”
Facilitator criterion: The participant reaches the request form and prepares the required sample information without route guidance. The endpoint is readiness to send, not a delivered enquiry. The facilitator checks the visible entries against the task card after the attempt; no sales or email integration is part of this task.
Task T3: interpret the next step
Read to the participant: “For this separate task, imagine the sample request has been sent. Starting from this simulated confirmation screen, tell us what has happened and what you would expect to happen next.”
Facilitator criterion: The participant distinguishes receipt of a quotation request from a confirmed booking. Any next step they describe must be supported by the screen. Reset everyone to this same simulated state, including anyone who could not complete T2. Record T3 as a separate attempt rather than evidence of an end-to-end successful submission.
Choose success criteria before sessions begin. Otherwise observers may reward the route they hoped to see or lower the standard after a difficult attempt. Allow equivalent successful routes. Finding accurate coverage through a search function can meet T1 just as well as navigating through the menu.
Pilot the script and control the starting conditions
Run a rehearsal before recruitment sessions. Ask a colleague unfamiliar with the script to attempt the tasks while the facilitator practises reading them and the observer takes notes. This is a check of the procedure. Label it as a rehearsal and keep it outside participant findings unless the study plan explicitly includes that person as a suitable participant.
Use the rehearsal to find broken prototype links, missing sample information and tasks with more than one interpretation. If T1 cannot be answered because the prototype contains no coverage policy, decide whether missing information is the intended question or a preparation mistake. Do not make participants spend the session finding defects you already know prevent every planned task.
Record a build identifier such as prototype-v1 and keep a copy of the task wording. Reset form values, navigation state and simulated confirmation screens between participants. When a device or connection changes, note the change. If an essential fix is made halfway through the round, label a new build and report attempts under their actual versions.
Task order also matters. T1 may teach someone where the company keeps information before T2 begins. That can be appropriate for a connected journey, but describe it. If your question concerns first impressions of two unrelated routes, consider varying their order and recording which came first. With a small sample, do not present order balancing as proof that learning effects have disappeared.
Harbour’s T3 deliberately starts at a supplied confirmation screen. This allows the team to examine its meaning even when T2 was difficult. It also creates a firm limit: completing T3 does not show that the participant successfully submitted a request. Keep that distinction in both the notes and the final report.
Facilitate without coaching the route
Use an opening that removes pressure and states the actual procedure. A notes-only version might read: “We are examining how this website works. You may pause, skip a task or stop. We will take the notes described in the information sheet; we will not record this session. Please use the supplied fictional details. Some screens are simulations, and nothing should be sent.”
After the consent process, introduce one task at a time. Let the participant read or hear it in the format that works for them. Invite them to describe what they are trying to do where comfortable, but do not insist on constant speech. A person may need to concentrate on reading, listening to assistive technology or planning an action.
When they ask, “Is this the right button?”, resist confirming your intended route. A neutral response is “What would you expect it to do?” If they are trying to remember the scenario, repeat the prompt without adding menu names. If your wording caused confusion, clarify the situation and record the clarification so it does not disappear from the evidence.
Silence is not automatically a problem. Someone may be comparing information carefully. Write down what you can observe before assigning a cause. When a pause matters, a prompt such as “What are you considering at this point?” is more informative than “Did you miss the button?” The latter suggests both a diagnosis and a preferred action.
Agree a rescue rule before the round. For this example, offer to move on when a participant says they cannot proceed, when continuing would exceed the planned session allowance or when a safety boundary is reached. If learning about a later step matters, supply the minimum help needed and record exactly what you said. Completion after route guidance is assisted completion. Repeating the unchanged prompt, providing an agreed interpreter or enabling someone’s usual interaction method is not automatically route guidance. Record these arrangements and clarifications; classify assistance by whether it revealed how to achieve the task.
A useful follow-up references an action: “You returned to the previous page twice. What were you looking for?” Avoid turning every hesitation into an interview about buying preferences. Keep questions tied to the task, the information encountered and the participant’s interpretation. Their explanation can clarify an observation without becoming proof of the underlying technical cause.
Brief observers too. They should not message hints, react visibly to mistakes or defend the design. If someone needs to ask a question, collect it for the facilitator. Close the session with a chance to add anything unclear, explain the next step and follow the agreed data-handling process.
Record actions, outcomes and interpretations separately
A compact observation log needs a stable participant code, task identifier, build, relevant context, outcome, assistance and factual notes. Add interpretations in a separate column. Link an issue identifier only after reviewing the evidence. Avoid writing a diagnosis into every row while the session is still unfolding.
For example, “F01 opened Services, returned to the homepage, then opened Contact” describes actions. “F01 expected coverage information near the enquiry route” is an interpretation unless supported by what they said. “Put a coverage summary beside the form” is a proposed change. Keeping these separate lets another reviewer challenge the explanation without discarding the observation.
GDS’s note-taking guidance recommends recording what is seen or heard and distinguishing it from interpretation. It also recommends labelling observations with session information. The outcome codes below are this guide’s own conventions; agree their meaning before using them. Read the guidance on research notes.
| Code | Meaning | Recording rule |
|---|---|---|
| complete | Agreed task criterion met without route guidance. | Record detours and mistakes even when the final outcome is correct. |
| assisted | Criterion met after help that revealed how to proceed. | Record the help; do not combine with independent completion. |
| incomplete | An attempt ended without meeting the criterion. | Describe the endpoint and observed obstacle without blaming the person. |
| stopped | A started attempt was interrupted by a boundary or external event. | Record why; do not automatically classify it as interface failure. |
| not_attempted | The task never began. | Keep it visible but exclude it from attempted-task counts. |
We test pages, forms and offers to turn more visitors into enquiries.
Use one row per participant-task attempt in the companion CSV. Summarise the important sequence in that row and use multiple issue IDs, separated by semicolons, if necessary. Do not create three attempt rows merely because someone encountered three problems. That would inflate the apparent number of opportunities to observe an issue. During analysis, agree a short description for each issue code, compare the linked observations and preserve contradictory evidence. If reviewers code a note differently, inspect the original note together and document the decision or uncertainty. Matching labels alone do not establish a shared cause.
Separate repetition within an attempt from the number of people who encountered a problem. Three unsuccessful taps by F01 remain one participant’s attempt. Record denominators by task and relevant exposure. If only three people saw the confirmation screen, the confirmation finding has three observed opportunities, even if four people attended sessions.
Task timing can add context if collected consistently, but notes-only sessions with interruptions are not precise performance benchmarks. Record an interruption or clarification rather than silently counting it as website effort. Do not calculate average completion time by treating incomplete tasks as zero or discarding them without explanation.
Work through a clearly fictional set of findings
Every entry in this example is invented. Harbour, the four participant codes, the prototype behaviour and the actions below do not describe an actual business or research study. The example demonstrates how to reason from a small evidence log without turning it into a population claim.
In T1, F01 finds the eligibility information after the facilitator points to the relevant page. F03 reaches the correct conclusion independently after two detours. F02 and F04 complete the task directly. The common issue is difficulty locating the policy, but the assistance and outcomes differ. The record should preserve that difference.
In T2, F01 cannot proceed because a mandatory unit-number field rejects a blank value for a standalone building. F03 proceeds after being told the fictional prototype accepts “N/A”. F02 and F04 complete the preparation task without help because their sample scenarios include unit numbers. This suggests investigating field logic, not claiming that all visitors need a shorter form.
In T3, F02 interprets the prominent simulated message “You’re all set” as a confirmed appointment and overlooks the less prominent explanation that only a request has been received. F01 and F03 correctly describe the intended next step. F04 does not attempt T3 because the fictional session ends early. The missing attempt remains in the dataset instead of becoming a failure or a success.
| Issue | Observed evidence | Supported next action |
|---|---|---|
| I01: coverage hard to locate | F01 and F03 encounter navigation difficulty in 2 of 4 T1 attempts; one completes with assistance. | Prototype a clearer information route and observe fresh attempts. |
| I02: irrelevant required unit field | F01 cannot complete T2; F03 completes with help. These are 2 of 4 T2 attempts, and both of the 2 attempts assigned the standalone-building scenario. | Confirm the business rule, then test a suitable optional or conditional field. |
| I03: receipt mistaken for booking | F02 misinterprets the message in 1 of 3 T3 attempts. One scheduled T3 task is unattempted. | Clarify receipt versus appointment status and examine comprehension again. |
I02 has two denominators worth keeping visible: all four T2 starts and the two standalone-building attempts exposed to the blank-unit condition. The CSV keeps four as the task denominator; its context column identifies the narrower exposure. Both standalone scenarios also use mobile devices, while the multi-unit scenarios use desktop. This constructed example cannot separate device effects from scenario effects.
The 12 scheduled task rows contain 11 started attempts: seven complete, two assisted and two incomplete. One task is unattempted. These totals are bookkeeping checks for the fictional file. They are not a website success-rate estimate, because they combine different tasks, deliberate sample scenarios and a tiny constructed group.
A defensible finding reads: “In this round, two participants had difficulty locating coverage information; one needed a directional hint.” An unsupported rewrite would be: “Half of customers cannot find our service area.” The first describes observations. The second invents a population estimate and overstates even the example’s task outcomes.
Keep alternative explanations alive. Maybe a task phrase caused people to search for terminology the site never uses. Maybe familiarity helped another participant. Maybe a prototype defect created the form restriction. Investigate these possibilities before announcing a root cause or committing to a larger redesign.
Prioritise issues with reasons, not a mysterious score
Severity describes the consequence of an issue. Priority describes when and how the team will act. A repair can be severe but require a temporary workaround while development is scheduled. A minor text change may be quick to release, but low effort does not make it more important than a blocked or misleading journey.
| Level | Consequence | Example interpretation |
|---|---|---|
| S3: critical task impact | Blocks an important task without viable independent recovery, or creates a materially wrong outcome. | A form cannot accept a legitimate scenario; a receipt is mistaken for a confirmed appointment. |
| S2: substantial friction | Creates meaningful detours, uncertainty or rework with a plausible recovery route. | Coverage information is discoverable but difficult to locate. |
| S1: limited friction | Causes a bounded inconvenience without changing the task’s outcome. | A label briefly delays an otherwise clear action. |
| S0: no demonstrated task issue | An opinion, suggestion or observation has no established adverse task effect. | A participant prefers a different accent colour. |
This rubric is an original aid to consistent discussion, not a scientific scale or a substitute for specialist assessment. Document the reason for each classification. Where the consequence is uncertain, mark it provisional and assign an investigation. Reviewers can disagree; the disagreement should produce a clearer question rather than an averaged number that hides it.
For Harbour, I02 and I03 receive S3 because they obstruct a legitimate request or create a mistaken booking expectation. I01 receives S2 for navigation friction with an available information route. All three deserve action, but the fictional team puts the form rule and confirmation wording first because their consequences extend beyond inconvenience.
Keep occurrence separate from severity. I03 appears once in the fictional file and can still be important. Conversely, a repeated preference for a different colour does not outweigh a serious barrier. Also distinguish observation confidence from cause confidence: a participant’s action may be clearly recorded while the explanation for it remains uncertain.
A useful issue ticket contains the task and build, affected attempt IDs, evidence, suspected explanation, consequence, proposed change, owner and a retest criterion. Give effort its own field. Do not multiply ordinal severity, invented audience percentages and guessed revenue into a number that appears more precise than the evidence.
Turn each issue into a repair and a retest condition
Write the smallest change that addresses the supported problem, then consider its wider effects. For I02, confirm whether the business actually needs a unit number for every building. If not, propose optional or conditional input. Do not simply remove the field from all journeys if operations need it for multi-unit buildings.
The acceptance criterion should describe behaviour: “The form permits a standalone-building sample without a unit number, while preserving the information needed for a multi-unit request.” A developer can check that rule with synthetic inputs. A participant retest can examine whether the form now makes sense without coaching. Those checks answer related but different questions.
For I03, first verify the real operational promise. If submitting a request does not reserve a visit, the confirmation should make that state clear. Do not invent a response deadline just to improve the wording. Ask the operations owner to approve any stated next step, then observe whether people understand it.
For I01, try an information route that makes coverage available where a visitor evaluates fit. Test whether participants locate and interpret it, not whether they click your preferred link. A new block might improve discovery while making an already long page harder to scan; observe the surrounding task as well as the changed element.
Use a new build identifier for the revised version. Fresh participants help examine whether a repair works without remembering the original path. Returning participants can provide useful feedback, but their prior exposure must remain visible. Retain task meaning and success criteria where possible, and record any changes that make rounds less comparable.
Do not call the repair a commercial success solely because the next small round looks easier. To assess enquiry quality or conversion impact, define the relevant outcomes and use an appropriate measurement design. Usability findings can justify a well-supported repair while its business effect remains unmeasured.
Use the task script and companion files
Download the usability-study templates (ZIP)
The companion set contains blank and fictional study plans in JSON, a facilitator script, blank and fictional observation and issue CSVs, and a field dictionary with instructions. Keep the fictional files as examples and start real work from the blank versions. The article’s task cards, outcome definitions and severity table also work as a copy-ready starting point.
Begin by completing the decision, audience, access arrangements, consent plan, build and tasks. Run the rehearsal. During sessions, use one observation row per task attempt and keep factual notes separate from interpretation. Afterwards, group evidence into issues and assign a specific repair owner and retest criterion.
The fictional files are checked for consistent task references, allowed outcome codes, issue links and the counts reported above. Those synthetic data checks do not constitute participant research, a browser study, form delivery testing, an accessibility audit or validation of a live integration. Your own test environment and organisational process still need examination.
When sharing findings, lead with the decision and a short evidence table. Include the recruited audience, number of started attempts per task, assistance, build, missing coverage and unresolved explanations. Preserve enough context for a colleague to understand the result without exposing participant identities. This produces a practical repair brief and an honest starting point for the next round.
