Purpose
Collaboratively define how we'll know the need is satisfied—specifying measurable and observable criteria spanning functional (task performance), emotional (user feelings), and social (perception/trust) dimensions—enabling objective A6-A7 validation testing rather than subjective "do users like it?" assessment.
Success criteria answer: "What does success look like? How will solutions be evaluated?" Without precise, multi-dimensional criteria, validation becomes opinion-based rather than need-centered, leading to debates about whether solutions "passed" and misalignment between what's built and what satisfies actual needs.
The Jobs-to-be-Done framework reveals needs have three dimensions:
- Functional: Task completion, performance, efficiency, accuracy
- Emotional: How users feel—confidence, anxiety reduction, satisfaction, control
- Social: How users are perceived—status, trust, credibility, collaboration
Comprehensive success criteria address all applicable dimensions, with explicit measurement methods for A6-A7 validation.
When to Use:
- After root cause analysis and boundary definition complete (Steps 3-4)—criteria must align with root causes being addressed
- Before A2 Ideation—solutions designed against criteria
- All validated needs advancing to A2—non-negotiable activity
- A6-A7 validation planning—criteria structure testing
When Can Be Abbreviated:
- Pure functional needs with obvious measurement (rare)—but still document emotional/social dimensions
- Exploratory/low-stakes innovations where formal validation not planned
Prerequisites:
- Completed root cause analysis (Step 3)
- Completed boundary definition (Step 4)
- A1.2 evidence accessible for reference
- Stakeholders available for 2-3 hour workshop
- Success Criteria Matrix template prepared
Complete Procedure
Preparation (Before Workshop)
1. Prepare Success Criteria Matrix Template
Create three-section matrix (functional, emotional, social):
| |p4cm|p2cm| ID | Criterion Definition | Measurement Method | Priority |
|---|---|---|---|
| 4|c|Functional Success Criteria | |||
| FC1 | [Task performance requirement] | [How measured in A6-A7] | M / N |
| FC2 | ... | ... | M / N |
| 4|c|Emotional Success Criteria | |||
| EC1 | [Emotional outcome] | [How measured in A6-A7] | M / N |
| EC2 | ... | ... | M / N |
| 4|c|Social Success Criteria | |||
| SC1 | [Social outcome] | [How measured in A6-A7] | M / N |
| SC2 | ... | ... | M / N |
2. Prepare Input Materials
- Need statement (from Step 1)
- Root cause analysis summary (from Step 3)
- Boundary definition (from Step 4)
- A1.2 key evidence excerpts (quotes about "what success would look like")
- Jobs-to-be-Done framework explainer (functional/emotional/social dimensions)
3. Schedule Workshop
- Duration: 2-3 hours
- Participants: 5-8 people itemize
- Essential: Product owner, user researcher, domain expert
- Recommended: Design lead, validation lead (who will conduct A6-A7)
- Optional: Engineering lead, business stakeholder itemize
- Send pre-read 2 days before: Need statement, root cause summary, JTBD framework explainer
Workshop Execution
SECTION 1: Context Setting (20 minutes)
[0-5 min] Review validated need and root causes
Brief recap:
- Validated need statement
- Root cause(s) identified
- Boundaries defined—which root causes addressing
- Reminder: Criteria must enable determining if solutions address root causes
[5-15 min] Introduce three-dimensional success criteria framework
Explain Jobs-to-be-Done dimensions with examples:
Functional criteria answer: "Can users accomplish the task?"
- Performance: Speed, accuracy, throughput
- Capability: Can do X that couldn't before
- Efficiency: Less effort, fewer steps, reduced error
- Example: "User completes task in <5 minutes" (measurable with time-to-task observation)
Emotional criteria answer: "How do users feel?"
- Confidence: Trust in their decisions/actions
- Anxiety reduction: Less stress, worry, overwhelm
- Satisfaction: Sense of accomplishment, progress
- Control: Agency over situation
- Example: "User reports feeling confident in decision" (measured with pre-post confidence survey)
Social criteria answer: "How are users perceived?"
- Status: Credibility, reputation, expertise demonstrated
- Trust: Others rely on user's output
- Collaboration: Strengthened relationships, teamwork
- Example: "Stakeholders express increased trust in user's forecasts" (measured through stakeholder feedback surveys)
Critical principle: Most needs have all three dimensions. Solutions satisfying only functional criteria while ignoring emotional/social needs will show low adoption in validation—users accomplish task but don't adopt because emotional/social needs unmet.
[15-20 min] Set workshop objectives
By end of workshop, team will have:
- 5-10 success criteria spanning all three dimensions
- Measurement method specified for each criterion
- Prioritization (must-have vs. nice-to-have)
- Evidence linkage (how criteria trace to A1.2)
SECTION 2: Functional Criteria Definition (30 minutes)
[20-30 min] Brainstorm functional criteria
Prompt: "If solutions address our root cause(s), what tasks should users be able to accomplish that they can't today? What performance improvements would indicate success?"
Process:
- Individual silent brainstorm (5 min): Write functional criteria on sticky notes, one per note
- Round-robin sharing (10 min): Participants share, facilitator places on matrix
- Clustering (5 min): Group similar criteria, identify duplicates
Functional criteria prompts:
- What specific tasks must users complete?
- How quickly must tasks be completed? (Time constraints)
- How accurately must tasks be completed? (Quality thresholds)
- What new capabilities must be enabled?
- What efficiency gains would indicate success? (Effort reduction, fewer steps, less rework)
[30-40 min] Refine and specify functional criteria
For each functional criterion brainstormed:
- Make specific: Avoid vague criteria itemize
- ✗ Poor: "Users can complete task easily"
- ✓ Good: "Users complete deal assessment task in ≤5 minutes without external tools" itemize
- Define measurement method: How will we measure this in A6-A7? itemize
- Task completion rate: % of users who successfully complete
- Time-to-task: Observed or logged time duration
- Accuracy: Comparison to ground truth or expert baseline
- Error rate: % of attempts with errors
- Efficiency: Steps required, tool-switching, workaround usage itemize
- Set threshold: What constitutes "passing"? itemize
- ✓ "≥80% of users complete task in ≤5 minutes"
- ✓ "≥90% accuracy compared to expert assessment" itemize
[40-50 min] Evidence linkage
For each functional criterion, ask: "What A1.2 evidence supports that users care about this?"
Cite interview quotes, observation behaviors, or workarounds that demonstrate users value this functional outcome.
If no evidence exists, criterion may be researcher projection—either find supporting evidence or mark as hypothesis to validate.
SECTION 3: Emotional Criteria Definition (30 minutes)
[50-60 min] Brainstorm emotional criteria
Prompt: "If solutions address our root cause(s), how should users feel differently? What emotional outcomes indicate need satisfaction?"
Emotional criteria prompts:
- What anxieties, frustrations, or stresses should be reduced?
- What confidence, satisfaction, or sense of control should be gained?
- What emotions did users express in A1.2 when describing need? (Review quotes)
- If need were satisfied, what would users say about how they feel?
Process: Same as functional (silent brainstorm, sharing, clustering)
[60-70 min] Refine and specify emotional criteria
For each emotional criterion:
- Make specific: Name the emotion precisely itemize
- ✗ Poor: "Users feel good about solution"
- ✓ Good: "Users express confidence in their assessments" OR "Users report reduced anxiety about forecast accuracy" itemize
- Define measurement method: itemize
- Pre-post surveys: Likert scales measuring confidence, anxiety, satisfaction, control itemize
- Example: "On scale 1-7, how confident are you in your deal assessment? [Before: X, After: Y]" itemize
- Qualitative interviews: Open-ended questions about feelings itemize
- "How did using this solution make you feel about [task]?" itemize
- Sentiment analysis: Coded language from think-aloud protocols or post-task interviews itemize
- Count positive emotional words (confident, comfortable, satisfied) vs. negative (anxious, frustrated, uncertain) itemize
- Behavioral proxies: Observable behaviors indicating emotional state itemize
- Hesitation, revisiting decisions, seeking confirmation from others suggests low confidence itemize itemize
- Set threshold: itemize
- ✓ "≥70% of users report increased confidence (≥+1 point on 7-point scale)"
- ✓ "≥60% of users use positive emotional language in post-task interview" itemize
[70-80 min] Evidence linkage
Cite A1.2 quotes revealing emotional dimensions:
- "I feel anxious every Sunday night before forecasting" → EC: Anxiety reduction
- "I wish I could feel more confident in my numbers" → EC: Confidence gain
- "It's stressful not knowing if I'm right" → EC: Stress reduction, certainty gain
SECTION 4: Social Criteria Definition (30 minutes)
[80-90 min] Brainstorm social criteria
Prompt: "If solutions address our root cause(s), how should users be perceived differently by others? What social/relational outcomes indicate success?"
Social criteria prompts:
- How do others (managers, peers, stakeholders) perceive user today regarding this need?
- If need satisfied, how should that perception change?
- What trust, credibility, or status shifts would indicate success?
- What collaboration or relationship improvements would occur?
- Did A1.2 reveal social consequences of unmet need? (Credibility damage, strained relationships, reduced influence)
Process: Same as functional/emotional (silent brainstorm, sharing, clustering)
[90-100 min] Refine and specify social criteria
For each social criterion:
- Make specific: Identify whose perception matters and what aspect itemize
- ✗ Poor: "Users are trusted more"
- ✓ Good: "VP expresses increased trust in manager's forecast accuracy" OR "Peers seek user's advice on [topic], indicating perceived expertise" itemize
- Define measurement method: itemize
- Stakeholder feedback surveys: Ask managers/peers about perception changes itemize
- "Has [user]'s forecast accuracy improved? Do you trust their forecasts more now?" itemize
- User perception surveys: Ask users if they feel others perceive them differently itemize
- "Do you feel your manager trusts your forecasts more? Do peers see you as more knowledgeable?" itemize
- Behavioral indicators: Observable relationship changes itemize
- Manager questions user's forecasts less frequently
- Peers ask user for advice more often
- User invited to strategic meetings (status elevation) itemize
- Collaboration metrics: Improved teamwork, communication, alignment itemize
- Cross-functional conflicts reduced
- Joint decision-making increased itemize itemize
- Set threshold: itemize
- ✓ "≥70% of managers report increased trust in forecasts"
- ✓ "≥50% of users report feeling more credible with stakeholders" itemize
[100-110 min] Evidence linkage
Cite A1.2 quotes revealing social dimensions:
- "My VP questions my numbers every week—I've lost credibility" → SC: Credibility restoration
- "I don't want to look incompetent" → SC: Competence perception
- "When I miss forecasts, the team loses confidence in me" → SC: Team trust
SECTION 5: Prioritization and Validation (30 minutes)
[110-125 min] Prioritize criteria (must-have vs. nice-to-have)
Review all criteria across three dimensions. Categorize:
Must-have (M): Non-negotiable for need satisfaction. If solution fails this criterion, need is NOT satisfied, regardless of other criteria.
Nice-to-have (N): Enhances satisfaction but not essential. Solutions can satisfy need without meeting this, though meeting it improves outcome.
Prioritization criteria:
- Evidence strength: Strongly evidenced in A1.2 → Must-have; Weak/inferred → Nice-to-have
- Root cause alignment: Directly addresses root cause → Must-have; Tangential → Nice-to-have
- User consensus: Universal across user types → Must-have; Specific to subset → Nice-to-have
- Impact: High impact on need satisfaction → Must-have; Marginal → Nice-to-have
Target: 3-5 must-have criteria total (not per dimension), 3-7 nice-to-have
Too many must-haves create overly-constrained solution space. Too few risk insufficient validation rigor.
[125-140 min] Validate completeness and consistency
Completeness check:
- Dimensional coverage: Do we have criteria in all three dimensions (functional, emotional, social)? itemize
- If missing dimension, revisit—most needs have all three
- Exception: Pure technical/process problems may lack strong social dimension itemize
- Root cause alignment: For each addressable root cause (from Step 3), do we have criteria indicating whether that cause is resolved? itemize
- If addressing "lack of health signals," criteria should include functional ability to assess using signals itemize
- Boundary alignment: Do criteria scope match boundaries (Step 4)? itemize
- If boundaries focus on Context X, criteria should specify Context X itemize
Consistency check:
- Do functional, emotional, social criteria align logically? itemize
- Example: If FC = "Assess deals in 5 min," EC = "Confidence," SC = "VP trust"—these align
- Inconsistency: If FC = "Assess deals quickly" but EC = "Reduced time pressure"—conflicting (speed may increase pressure) itemize
- Are thresholds realistic? itemize
- 80-90% success rates reasonable for must-have functional
- 60-70% for emotional (harder to measure precisely)
- 50-70% for social (depends on stakeholder feedback availability) itemize
Post-Workshop Documentation
Within 24 hours:
- Finalize Success Criteria Matrix (table format) itemize
- All criteria with IDs, definitions, measurement methods, priorities
- Evidence citations for each criterion
- Thresholds specified itemize
- Create Measurement Protocol Document (2-3 pages) itemize
- For each criterion, detailed measurement protocol
- Survey instruments (specific questions, scales)
- Observation protocols (what to observe, how to code)
- Interview guides (questions to ask in A6-A7 validation interviews)
- Data collection and analysis methods itemize
- Circulate for Validation itemize
- Send to workshop participants for final review
- Send to A6-A7 validation team for feasibility check (Can we measure these?)
- Incorporate feedback and finalize itemize
- Incorporate into Problem Definition Brief itemize
- Success Criteria Matrix becomes Section 6 of brief
- Measurement protocols become appendix or separate validation planning document itemize
Quality Criteria
1. Dimensional comprehensiveness: Criteria span functional, emotional, and social dimensions (not just functional)
2. Specificity: Each criterion is precise and measurable/observable
- ✗ Poor: "Users are satisfied"
- ✓ Good: "Users report ≥6/7 satisfaction with assessment process"
3. Measurement methods defined: Every criterion has explicit measurement approach for A6-A7
4. Evidence-based: Criteria trace to A1.2 findings—not researcher preferences
5. Root cause-aligned: Criteria enable determining if root causes addressed
6. Boundary-aligned: Criteria scope matches boundaries (user types, contexts, functions)
7. Prioritized: Must-have vs. nice-to-have distinction clear (3-5 must-have total)
8. Validated: Stakeholders agree these criteria capture need satisfaction
9. Feasible to measure: A6-A7 validation team confirms they can measure criteria
Tools and Resources
Physical workshop:
- Large wall space or whiteboards (3 sections for functional/emotional/social)
- Sticky notes (3 colors—one per dimension)
- Markers for annotations
- Camera for documentation
Digital workshop:
- Miro, Mural, or FigJam with Success Criteria Matrix template
- Video conferencing
- Collaborative editing enabled
Templates:
- Success Criteria Matrix (Excel, Google Sheets, Miro template)
- JTBD framework explainer (functional/emotional/social dimensions)
- Measurement method selector (decision tree for choosing measurement approach)
- Validation protocol template (for A6-A7 planning)
Sample Size / Duration
Participants: 5-8 people
- Essential: Product owner, user researcher (conducted A1.2), domain expert
- Recommended: Validation lead (will conduct A6-A7), design lead
- Optional: Engineering lead, business stakeholder
- Avoid: >10 people
Duration:
- Workshop: 2-3 hours (20 min context + 30 min per dimension + 30 min prioritization)
- Pre-work: 30 minutes (reviewing need, root causes, JTBD framework)
- Post-workshop documentation: 3-4 hours (finalizing matrix, measurement protocols)
- Total: 6-8 hours (primarily facilitator and product owner time)
Common Challenges and Solutions
Challenge 1: Emotional/social dimensions ignored—only functional criteria
Symptoms:
- Matrix has 8 functional criteria, 0-1 emotional, 0 social
- Team says "emotional/social are too subjective to measure"
- Criteria focus on task performance, ignore user feelings/perception
Solutions:
- JTBD framework enforcement: Facilitator mandates "We must have at least 2 emotional and 2 social criteria before moving on"
- Evidence review: Return to A1.2 quotes—highlight emotional language ("I feel anxious," "I worry about...," "I don't want to look incompetent")—these reveal emotional/social dimensions
- Low adoption consequence: Explain that solutions satisfying only functional needs show low adoption because emotional/social needs unmet
- Measurement examples: Provide concrete measurement methods for emotional/social to reduce "too subjective" concern
Challenge 2: Vague, unmeasurable criteria
Symptoms:
- Criteria like "easy to use," "intuitive," "better experience"
- No measurement method specified—how would we measure this?
- Criteria couldn't objectively determine pass/fail in validation
Solutions:
- Specificity test: For each criterion, ask "How would we measure this in A6-A7? What would we observe or ask?" itemize
- If team can't answer, criterion needs refinement itemize
- Operationalization: Transform vague criteria into measurable proxies itemize
- "Easy to use" → "≥80% users complete task in ≤5 min without help"
- "Intuitive" → "≥70% users don't need external documentation"
- "Better experience" → "Users rate satisfaction ≥6/7 on post-task survey" itemize
- Concrete behaviors: Ask "What would users do or say if this criterion is satisfied?"
Challenge 3: Too many criteria (>15)—validation becomes overwhelming
Symptoms:
- Matrix has 20+ criteria
- Everything feels important—can't prioritize
- A6-A7 validation team says "We can't feasibly measure all these"
Solutions:
- Consolidation: Combine overlapping criteria itemize
- "Complete task in 5 min" + "Complete task without external tools" → "Complete task in ≤5 min without external tools" itemize
- Prioritization forcing: Use dot voting—each participant gets 5 votes, allocates to top criteria
- Must-have limit: Enforce max 5 must-have criteria—forces hard choices
- Validation feasibility check: A6-A7 team reviews criteria—which are feasible to measure? Deprioritize infeasible ones
Challenge 4: Criteria don't align with root causes
Symptoms:
- Root cause is "Lack of health framework" but criteria focus on "Forecast accuracy"—misalignment
- Criteria measure outcomes several steps removed from root cause
- Solutions could address root cause but fail criteria (or vice versa)
Solutions:
- Traceability check: For each addressable root cause, identify at least 1-2 criteria directly indicating if that cause is resolved itemize
- Root cause: "No health framework" → Criterion: "User assesses deals using standardized health signals" itemize
- Causal logic: Validate that satisfying criteria logically follows from addressing root cause itemize
- If health framework provided, THEN users can assess using signals (logical)
- If health framework provided, THEN forecasts are accurate (indirect—many other factors) itemize
- Proximate criteria: Favor criteria directly reflecting root cause resolution over distal outcomes
Example: Sales Forecast Confidence Criteria
Context: Root cause is "Organization hasn't defined deal health framework; CRM shows stage not health signals"
Success Criteria Matrix:
| |p4cm|c| ID | Criterion | Measurement (A6-A7) | Priority |
|---|---|---|---|
| 4|c|Functional Success Criteria | |||
| FC1 | Manager assesses any deal's closure probability within 5 minutes using evidence (not gut-feel) | Time-to-task observation; evidence usage check (uses signals not gut) | M |
| FC2 | Manager can identify 3-5 specific health signals supporting their assessment | Post-task interview: "What signals did you use?" Count signals cited | M |
| FC3 | Assessment accuracy improves vs. baseline (≥15 percentage points) | Compare manager forecast accuracy to baseline over 4 weeks | N |
| 4|c|Emotional Success Criteria | |||
| EC1 | Manager expresses increased confidence in forecast | Pre-post confidence survey: "How confident in forecast (1-7)?" Target ≥+1 point improvement, ≥70% managers | M |
| EC2 | Manager reports reduced Sunday evening anxiety | Weekly diary: "How anxious about forecast (1-7)?" Target ≥-1 point reduction | N |
| 4|c|Social Success Criteria | |||
| SC1 | VP expresses increased trust in manager's forecasts | VP stakeholder survey: "Do you trust [manager]'s forecasts more? (Y/N)" Target ≥70% Yes | M |
| SC2 | Peers seek manager's advice on deal assessment (perceived expertise elevation) | Manager self-report: "Have peers asked your advice on deals? (Y/N)" Target ≥50% Yes | N |
Measurement Protocol Notes:
- FC1: A6 usability testing with 10-12 managers, measure time and observe whether they cite data or use gut-feel language
- FC2: Post-task interview immediately after assessment task
- FC3: A7 pilot with 15-20 managers over 4 weeks, compare forecast accuracy to 4-week pre-pilot baseline
- EC1: Pre-post survey administered at start of A6/A7 and after 2 weeks usage
- EC2: Weekly diary during A7 pilot (opt-in 8-10 managers)
- SC1: Survey VPs after A7 pilot (4 weeks)—one survey per manager's VP
- SC2: Manager survey at end of A7 pilot
Must-have rationale:
- FC1, FC2: Directly indicate root cause addressed (using health signals, not gut-feel)
- EC1: Core emotional need from A1.2—confidence was primary emotional pain point
- SC1: Social dimension critical per A1.2—credibility with VP drives behavior
Nice-to-have rationale:
- FC3: Accuracy improvement is ideal outcome but indirect—many factors beyond health signals affect accuracy (market conditions, deal complexity)
- EC2: Anxiety reduction valuable but confidence (EC1) more directly tied to root cause
- SC2: Peer perception enhancement valuable but not essential—VP trust (SC1) more critical
Share how you use Success Criteria Workshop
This is where practitioners will be able to share field notes, variations, and additional templates for this method — what worked, what to watch for, and adaptations for different contexts.
Until the community space opens, we welcome contributions by email and will fold the best into the method page.