Learning Resources · Methods Library · System Usability Scale (SUS)
IdeationDeliveryScaling

System Usability Scale (SUS)

Used in: A3.2 Steps 3–5 (post-session satisfaction measurement)

Also applicable: A3.3 (pilot satisfaction tracking), A2.3 (lo-fi baseline), A4–A7 (post-launch benchmarking)

QR code linking to this method page Scan to open

Purpose

Provide a quick, reliable, and standardised measure of perceived usability after users interact with a prototype. SUS produces a single composite score (0–100) that enables comparison against industry benchmarks, across prototypes, and across testing waves—answering “How usable do users perceive this to be?” as a complement to the behavioural metrics (task completion, time-on-task) captured by task-based usability testing (the referenced method).

SUS is the most widely used standardised usability questionnaire, with over 1,300 published studies providing robust benchmarking data. Its 10-item format takes under 2~minutes to complete, making it practical to administer at the end of every A3.2 testing session without fatiguing participants.

When to Use

Use SUS when:

  • After A3.2 testing sessions—administered immediately post-tasks while the experience is fresh
  • At regular intervals during A3.3 pilot (Weeks~2, 4, 6, 8) to track perceived usability over time
  • Comparing two prototype variants (SUS scores compared using t-test or Mann-Whitney U)
  • Benchmarking against industry averages or competitor products
  • Tracking improvement across A3.2 Wave~1 Wave~2 (did iteration improve perceived usability?)

Do NOT use when:

  • As the sole usability measure—SUS captures perception, not performance; a user may rate a system highly but fail tasks, or rate it poorly but succeed
  • Before any interaction—SUS measures experience, not expectation
  • With fewer than 12 respondents—SUS scores become unreliable below this threshold
  • To diagnose specific problems—SUS identifies overall satisfaction level but not which features cause problems (use task-based testing and think-aloud for diagnosis)

Sample Size and Duration

Respondents: 12 minimum; 20–50 recommended (administered to all A3.2 test participants)

Administration time: 2 minutes per participant

Scoring and analysis: 30–60 minutes per prototype (once all responses collected)

Prerequisites

  • User interaction: Participants must have used the prototype (minimum: completed core tasks in testing session)
  • Standardised form: The 10 items must be presented exactly as written—rewording invalidates the benchmarking data
  • Timing: Administered immediately after task completion, before any debrief discussion
  • Sample: ≥12 respondents for reliable mean score; ≥20 for comparisons between groups

Complete Procedure

Step~1: Administer SUS (2 minutes per participant)

Present the 10-item questionnaire immediately after the participant completes all test tasks. Paper form, digital survey (Google Forms, Typeform), or embedded in testing platform. Instruct: “Please rate your immediate response—don't think too long about each item.”

Step~2: Score Individual Responses (5 minutes per batch)

Scoring follows a specific algorithm:

  1. For odd-numbered items (1, 3, 5, 7, 9): subtract 1 from the user's response (score contribution = response - 1)
  2. For even-numbered items (2, 4, 6, 8, 10): subtract the user's response from 5 (score contribution = 5 - response)
  3. Sum all 10 score contributions (range: 0–40)
  4. Multiply by 2.5 to produce the SUS score (range: 0–100)

Important: SUS scores are not percentages. A score of 68 does not mean “68% usable”—it means “average usability” based on benchmark data.

Step~3: Compute Aggregate Statistics (15 minutes)

  • Mean SUS score across all participants (primary metric)
  • Standard deviation (variability indicator)
  • 95% confidence interval (precision of estimate)
  • Per-variant means if A/B testing (the referenced method)

Step~4: Interpret Against Benchmarks (15 minutes)

p2.5cmp3cmp5.5cm SUS ScoreGradeAdjectiveA3.2 Interpretation
≥80.3AExcellentStrong pass; ready for A3.3
68–80.2B–CGood–OKPass; iteration may improve
51–67.9DPoorMarginal; significant iteration needed
≤50.9FAwfulFail; fundamental redesign required
SUS score interpretation for A3.2 decisionssauro2016quantifying

A3.2 target: ≥68 (above average). The running example's success criteria specify SUS ≥68 as the usability threshold for A3.3 pilot advancement.

Step~5: Report (included in validation report)

Report: mean SUS, confidence interval, benchmark comparison, per-variant comparison (if applicable), Wave~1 vs. Wave~2 improvement (if applicable). Include individual item analysis if specific dimensions need investigation (e.g. items 2 and 8 relate to complexity; items 4 and 10 relate to learnability).

Quality Criteria

  1. Standardised administration: 10 items presented exactly as written; no modifications
  2. Correct scoring: Algorithm applied correctly (odd/even reversal, × 2.5)
  3. Adequate sample: ≥12 respondents; ≥20 for comparisons
  4. Benchmark-contextualised: Score reported with industry comparison, not as standalone number
  5. Confidence interval reported: Precision of estimate visible to decision-makers
  6. Complemented by behaviour: SUS presented alongside task completion and time-on-task (perception + performance)

Theoretical Foundation

Seminal references:

  • : Created SUS as a “quick and dirty” usability scale. Despite the modest label, subsequent validation studies confirmed excellent psychometric properties: high reliability (Cronbach's = 0.91), sensitivity to differences between systems, and concurrent validity with other usability measures.
  • : Compiled SUS benchmark data from 500+ studies (5,000+ users), establishing that the average SUS score is 68. Scores above 68 are above average; below 68 are below average. Provided percentile rankings and letter-grade interpretations that make SUS scores actionable for non-researchers.

Contemporary references:

  • : Validated the adjective-anchored interpretation scale (Worst Imaginable Best Imaginable) and confirmed SUS's reliability across product categories, user demographics, and testing modalities (lab vs. remote).

The 10 SUS Items

SUS consists of 10 statements rated on a 5-point Likert scale (Strongly Disagree to Strongly Agree). Odd-numbered items are positively worded; even-numbered items are negatively worded (reducing acquiescence bias):

  1. I think that I would like to use this system frequently.
  2. I found the system unnecessarily complex.
  3. I thought the system was easy to use.
  4. I think that I would need the support of a technical person to be able to use this system.
  5. I found the various functions in this system were well integrated.
  6. I thought there was too much inconsistency in this system.
  7. I would imagine that most people would learn to use this system very quickly.
  8. I found the system very cumbersome to use.
  9. I felt very confident using the system.
  10. I needed to learn a lot of things before I could get going with this system.

Challenges and Solutions

Challenge 1: Score Misinterpretation

Symptoms: Team interprets SUS 72 as “72% usable” or “C grade” and panics.

Solutions: Always present SUS with benchmark context: “72 is above the industry average of 68, placing this prototype in the 55th percentile. This is a B- grade—good, with room for improvement.”

Challenge 2: Small Sample Instability

Symptoms: SUS from 8 users is 74; team celebrates. Actual confidence interval: 58–90.

Solutions: Report confidence intervals. With 8 users, the CI is too wide for reliable conclusions. Minimum 12 for stable estimates; 20+ for meaningful comparisons.

Challenge 3: Ceiling Effect in Controlled Testing

Symptoms: SUS is 82 in A3.2 testing (facilitator present, quiet lab) but drops to 65 in A3.3 pilot (real-world distractions, self-service).

Solutions: Expect 5–10 point drop between controlled testing and real-world usage. If A3.2 SUS is borderline (68–72), it may fall below threshold in A3.3—flag this risk in the validation report.

Relationship to Other Methods

SUS receives input from:

  • Task-Based Usability Testing (the referenced method): SUS administered immediately after tasks; task experience shapes SUS response

SUS provides input to:

  • A3.4 Evaluation: SUS score is a key desirability-lens metric
  • A3.3 Pilot: Longitudinal SUS tracking (Weeks~2, 4, 6, 8) monitors perceived usability over sustained use
  • A/B Testing (the referenced method): SUS comparison between variants provides secondary decision metric

SUS is complemented by:

  • NPS (the referenced method SUS measures usability perception; NPS measures loyalty/recommendation intent—different constructs
  • Think-Aloud Protocol (the referenced method): SUS identifies the overall satisfaction level; think-aloud reveals why users feel that way

Tools and Templates

  • Google Forms / Typeform / Qualtrics (survey administration)
  • SUS scoring spreadsheet template (Google Sheets / Excel with built-in scoring formula)
  • MeasuringU SUS calculator (online scoring and benchmarking tool)
  • Maze / UserTesting (built-in SUS administration in testing platforms)
  • A. Bangor, P. T. Kortum & J. T. Miller (2009). Determining What Individual SUS Scores Mean: Adding an Adjective Rating Scale. Journal of Usability Studies. 4(3). pp. 114–123.
  • J. Brooke (1996). SUS: A “Quick and Dirty” Usability Scale. Usability Evaluation in Industry. pp. 189–194.
  • J. Sauro & J. R. Lewis (2016). Quantifying the User Experience: Practical Statistics for User Research. 2 ed. Morgan Kaufmann.
Coming soon

Share how you use System Usability Scale (SUS)

This is where practitioners will be able to share field notes, variations, and additional templates for this method — what worked, what to watch for, and adaptations for different contexts.

Until the community space opens, we welcome contributions by email and will fold the best into the method page.