Learning Resources · Methods Library · Affinity Diagramming
Discovery

Affinity Diagramming

Used in: A1.2 Step 4 (Need synthesis) A1.3 Step 3 (Root cause synthesis) A6 Step 5 (Feedback synthesis) Related Activities: Any activity requiring thematic organization of qualitative data

QR code linking to this method page Scan to open

Purpose

Systematically organize large volumes of qualitative data (interview transcripts, observation notes, user feedback, diary entries, stories) into coherent patterns and themes through collaborative spatial clustering.

For example, after conducting 12-20 interviews and 4-8 observations, teams face synthesis challenge: hundreds of data points need organization to reveal meaningful patterns. Affinity diagramming provides systematic method for consolidating findings into thematic groups.

When to Use

  • 150+ discrete data points (insights, quotes, observations) requiring organization
  • Team-based synthesis needed (multiple researchers or stakeholders involved in research)
  • Pattern identification from unstructured qualitative data
  • Need to surface themes that aren't predetermined (emergent analysis)

When NOT to Use:

  • Small datasets (<100 data points)—manual review sufficient
  • Purely quantitative data—use statistical methods instead
  • Predetermined category systems—use deductive coding frameworks
  • Individual analyst work—affinity diagramming designed for collaborative synthesis

Prerequisites:

  • Completed data collection (interviews, observations, etc.)
  • Transcripts/notes prepared and reviewed
  • Atomic insights extracted (Step 2.6 Data Processing in A1.2)
  • Team availability for 4-8 hour synthesis workshop
  • Physical wall space + sticky notes OR digital whiteboard (Miro/Mural) + remote access

Complete Procedure

Preparation (before workshop)

1. Extract atomic insights from all research data

Review interview transcripts, observation field notes, diary study data, and story collections. Extract discrete insights onto individual sticky notes or digital cards (Miro, Mural). One insight per note.

Types of insights to capture:

  • Need statements: "Users struggle to [action] when [context]"
  • Obstacles: "Cannot accomplish [goal] because [barrier]"
  • Workarounds: "Currently cope by [approach]"
  • Emotional responses: "Feel [emotion] when [situation]"
  • Contextual factors: "Need intensity increases when [condition]"
  • User quotes: Direct verbatim statements revealing needs

Good practice:

  • Write in user's voice when possible: "I can't see which deals are actually likely to close"
  • Keep atomic—one idea per note, not complex multi-part statements
  • Attribute source—note which interview/user (enables tracking back, understanding user type patterns)
  • Target 200-400 individual notes for comprehensive research dataset

2. Prepare workspace

Physical approach:

  • Large wall space (10-15 feet wide minimum)
  • Sticky notes in single color for insights
  • Different colored notes for cluster labels (helps visual distinction)
  • Markers for labeling

Digital approach (Miro, Mural, FigJam):

  • Create large canvas/board
  • Upload all insights as individual cards
  • Prepare separate label cards in different color
  • Test that all team members have edit access

Workshop Execution (4-8 hours)

Team composition: All researchers who conducted interviews/observations should participate. Direct research exposure essential for interpreting insights—delegating synthesis to non-researchers loses nuanced understanding.

Facilitation: Appoint facilitator to manage process, timing, ensure all voices heard, probe for deeper patterns, document decisions.

Step 1: Silent clustering (60-90 minutes)

Spread all insight notes on wall or digital whiteboard. Team works silently to cluster related insights.

Process:

  1. Read through all notes individually (15-20 min)
  2. Begin moving notes into spatial clusters—identify similar notes describing related needs, obstacles, contexts (40-60 min)
  3. Continue until initial small clusters form (3-7 related notes per cluster)
  4. No discussion during this phase—prevents groupthink, allows multiple perspectives to emerge

Clustering principle: Bottom-up, emergent. Let patterns emerge from data rather than forcing into preconceived categories. No predetermined framework.

Result: Many small clusters initially (30-50 clusters typical for 300 notes).

Step 2: Cluster review and consolidation (30-45 minutes)

Team discusses clustering collectively:

  • Review each cluster: Do these notes belong together? What connects them?
  • Consolidate overlapping clusters if multiple people created similar groupings
  • Split clusters if grouping too heterogeneous (combining unrelated insights)
  • Move orphan notes that don't fit current clusters

Decision rule: Notes within cluster should describe related aspect of same need theme. If debate whether note belongs, try both placements and see which creates more coherent cluster.

Step 3: Label clusters (45-60 minutes)

For each cluster, collaboratively create descriptive label capturing theme essence.

Labeling guidelines:

  • Use user-centric language reflecting user experience, not abstract categorization
  • Be specific: "Forecast confidence anxiety" (specific) vs. "Forecasting issues" (generic)
  • Capture core need/obstacle: Label should communicate what cluster is about without reading notes
  • 3-6 words optimal: Concise but meaningful

Examples:

  • "Data integration challenges across systems"
  • "Time pressure in quarterly reporting"
  • "Lack of deal progression visibility"
  • "Manual workaround burden"

Write label on distinctly colored note, place above cluster.

Step 4: Hierarchical grouping (60-90 minutes)

Look for patterns across clusters—do multiple clusters relate to broader theme?

Process:

  1. Arrange related clusters spatially near each other
  2. Create higher-level theme labels for cluster groups
  3. Typically creates 2-3 levels of hierarchy: itemize
  4. Level 1: Individual insights (200-400 notes)
  5. Level 2: Clusters (30-50 thematic groups)
  6. Level 3: Super-clusters (8-15 major need categories) itemize

Example hierarchy:

  • Super-cluster: "Data infrastructure needs" itemize
  • Cluster: "Data access barriers"
  • Cluster: "Data quality concerns"
  • Cluster: "Data integration challenges"
  • Cluster: "Real-time data availability" itemize

Step 5: Cross-cutting pattern identification (45-60 minutes)

Beyond hierarchical organization, identify patterns cutting across clusters:

User type patterns:

  • Do certain user types (if identified) contribute disproportionately to specific clusters?
  • Mark clusters with user type prevalence: "Primarily Type A" or "All user types"
  • Reveals user type-specific needs vs. universal needs

Context patterns:

  • Do certain contexts (time pressure, resource constraints, organizational factors) appear across multiple clusters?
  • Reveals contextual amplifiers—factors that intensify multiple needs simultaneously

Intensity patterns:

  • Which clusters contain highest emotional language? ("frustrated," "anxious," "overwhelmed")
  • Which clusters show most elaborate workarounds? (High effort indicates high need significance)
  • Mark high-intensity clusters—signals priority areas

Outputs

  • Visual synthesis diagram: Large-format diagram showing all insights organized into thematic clusters and hierarchies (photograph if physical, export if digital)
  • Theme documentation: Written description of each major theme (8-15 super-clusters) including: itemize
  • Theme label and description
  • Representative quotes from cluster
  • User type associations (which types experience this theme most)
  • Intensity indicators (emotional language, workaround burden)
  • Prevalence (how many users contributed to this cluster) itemize
  • Pattern insights: Cross-cutting patterns identified (user type patterns, contextual factors, intensity rankings)
  • Need prioritization: Clusters ranked by combination of prevalence, intensity, workaround burden, and strategic fit

Complete Procedure

Preparation (Before Workshop)

1. Prepare Canvas Template

Create four-quadrant canvas (physical poster 36"×24" or digital Miro/Mural board):

|p5.5cm| QUADRANT 1: User/Stakeholder BoundariesQUADRANT 2: Functional Boundaries
WHO experiences need? WHO benefits?WHAT functions/jobs will solution address?
• In-scope user types:• In-scope functions:
• Out-of-scope user types:• Out-of-scope functions:
• Expansion criteria:• Expansion criteria:
QUADRANT 3: Contextual BoundariesQUADRANT 4: Root Cause Boundaries
WHERE/WHEN does solution apply?WHICH root causes will we address?
• In-scope contexts:• Addressable causes (in-scope):
• Out-of-scope contexts:• Constraint causes (accept as given):
• Expansion criteria:• Deferred causes (future):
• Expansion criteria:

2. Prepare Input Materials

  • Print/display validated need statement
  • Summary of root cause analysis (causes identified, evidence, addressability assessment)
  • User type segmentation (if completed in A1.2)
  • List of stakeholders who must attend workshop (product owner, tech lead, business sponsor)

3. Schedule Workshop

  • Duration: 90-120 minutes
  • Participants: 5-8 people maximum (decision-makers, not observers)
  • Required attendees: product owner, technical lead, business stakeholder
  • Send pre-read 2-3 days before: need statement, root cause summary (ensures informed discussion)

Workshop Execution

SECTION 1: Context Setting (15 minutes)

[0-5 min] Review validated need and root causes

Present brief recap (5 slides maximum):

  • Validated need statement
  • User types experiencing need (if segmented)
  • Root cause analysis summary—key causes identified
  • Evidence snapshot (3-4 compelling quotes/observations)

[5-15 min] Introduce boundary canvas framework

Explain four quadrants:

  • Quadrant 1 (Users/Stakeholders): Which user types will solution serve? Initial release vs. future expansion?
  • Quadrant 2 (Functional): Which jobs/functions will solution address? Core vs. adjacent functions?
  • Quadrant 3 (Contextual): Which contexts/circumstances? Always applicable or context-specific?
  • Quadrant 4 (Root Causes): Which causes will we address vs. accept as constraints vs. defer?

Key principle: Boundaries are strategic choices, not arbitrary limits. Each boundary has trade-off—narrower enables focus and quality, broader serves more users but risks dilution.

SECTION 2: Define User/Stakeholder Boundaries (20 minutes)

[15-25 min] Brainstorm user types and stakeholders

Prompt: "Who experiences this need? Who else is affected by it?"

If A1.2 identified user types (2×2 segmentation), start there. Otherwise, brainstorm:

  • Primary users (directly experience need)
  • Secondary users (affected by need but not primary experiencer)
  • Stakeholders (don't experience need but have interest: managers, admins, compliance)

[25-35 min] Prioritize and define boundaries

For each user type/stakeholder, assess:

  • Need intensity: How severely does this group experience need?
  • Frequency: How often?
  • Strategic value: Alignment with business priorities?
  • Solution complexity: Does serving this group add disproportionate complexity?

Decision framework:

  • In-scope: High intensity + high frequency + strategic value + manageable complexity
  • Out-of-scope: Low intensity OR low frequency OR low strategic value OR excessive complexity
  • Expansion criteria: "Add [user type] when [condition]"—e.g., "Add enterprise sales managers when SMB version validated and revenue >$1M"

Example (Sales Forecast Confidence):

  • In-scope: Mid-market B2B sales managers (7-25 reps)
  • Out-of-scope: Enterprise sales VPs (>100 reps—different workflow), Solo sellers (<5 reps—different tools)
  • Expansion: "Add enterprise VPs when mid-market achieves 80% weekly active usage"

SECTION 3: Define Functional Boundaries (20 minutes)

[35-45 min] Map job-to-be-done workflow

Starting from validated need, map complete workflow where need arises:

  • What is user trying to accomplish? (Goal)
  • What steps involved? (Workflow stages)
  • Where does validated need arise? (Specific stage)
  • What adjacent functions exist? (Before/after stages)

Example (Sales Forecast):

  • Complete workflow: Data gathering → Deal assessment → Forecast aggregation → Executive reporting → Retrospective analysis
  • Validated need location: "Deal assessment" stage (understanding deal health/risk)
  • Adjacent functions: Forecast aggregation, reporting

[45-55 min] Define functional scope

Decide which workflow stages solution will address:

  • Core function: Directly addresses validated need (mandatory)
  • Adjacent functions: Closely related jobs that enhance core (optional but valuable)
  • Distant functions: Related workflow but not directly connected to validated need (out-of-scope)

Decision heuristics:

  • Include adjacent function IF: Solves core need requires it (dependency) OR significantly amplifies core value (10x better vs. 20% better)
  • Exclude distant function IF: Can address core without it AND adds substantial complexity

Example:

  • In-scope: Deal health assessment (core), deal risk signaling (adjacent—amplifies core)
  • Out-of-scope: Forecast aggregation (separate tool handles), executive reporting (different user), retrospective analysis (separate workflow)
  • Expansion: "Add forecast aggregation when deal assessment achieves 70% manager adoption"

SECTION 4: Define Contextual Boundaries (20 minutes)

[55-65 min] Identify contexts where need arises

From A1.2 discovery, recall contexts where need manifests:

  • Temporal: When during day/week/month? Time-sensitive moments?
  • Environmental: Where—office, home, mobile, specific locations?
  • Circumstantial: What conditions amplify need—time pressure, resource constraints, high stakes?
  • Social: Alone or collaborative? Individual work or team coordination?

[65-75 min] Define contextual scope

Not all contexts equally important or feasible to address:

Assessment criteria:

  • Need intensity by context: Which contexts show highest severity?
  • Frequency by context: Which contexts most common?
  • Solution feasibility: Which contexts can solution realistically serve? (Mobile context requires mobile design—significant investment)

Example:

  • In-scope: Weekly forecast preparation (Sunday evening/Monday morning), office desktop environment, individual analysis phase
  • Out-of-scope: In-meeting forecast defense (real-time, high pressure—different solution requirements), mobile/on-the-go (requires mobile app—deferred), collaborative forecast sessions (different workflow)
  • Expansion: "Add in-meeting support when individual analysis achieves 80% accuracy"

SECTION 5: Define Root Cause Boundaries (25 minutes)

[75-90 min] Review root cause analysis

Display root causes identified in Step 3 (from Five Whys, How-Why Laddering, or Ishikawa):

For each cause, assess addressability:

  • Can we change this? (Technical/organizational feasibility)
  • Should we change this? (Strategic fit, resource availability)
  • What leverage does addressing this provide? (Impact on need)

[90-100 min] Categorize causes into three buckets

1. Addressable causes (in-scope):

  • Within team capability to solve
  • High leverage (addressing significantly alleviates need)
  • Aligned with strategic priorities
  • Example: "Lack of structured deal health signals" ✓ Can build signals, high impact

2. Constraint causes (accept as given):

  • Outside control (regulatory, market forces, organizational culture)
  • Politically/technically infeasible to change
  • Example: "Sales compensation structure rewards sandbagging" → Compensation outside product team scope, accept as constraint

3. Deferred causes (future):

  • Addressable in principle BUT requires capabilities not yet developed
  • Lower leverage than in-scope causes—prioritize high-leverage first
  • Example: "Inconsistent CRM data entry" → Requires data quality infrastructure, defer to V2

Boundary principle: Solution should address 2-4 high-leverage causes comprehensively rather than 10+ causes superficially.

SECTION 6: Document Expansion Criteria (10 minutes)

[100-110 min] Define triggers for boundary expansion

For each "out-of-scope" or "deferred" element, specify conditions under which scope would expand:

Expansion trigger types:

  • Validation-based: "Add [scope element] when [adoption metric] reaches [threshold]"—e.g., "Add enterprise users when SMB achieves 75% retention"
  • Capability-based: "Add [scope element] when [technical capability] developed"—e.g., "Add mobile support when real-time sync infrastructure built"
  • Strategic-based: "Add [scope element] when [business condition]"—e.g., "Add forecast reporting when executive sponsors prioritize"

Purpose: Expansion criteria prevent scope debates ("why aren't we doing X?") by documenting decisions and future conditions.

Post-Workshop Documentation

Within 24 hours:

  • Photograph canvas or export digital whiteboard
  • Create boundary definition document (2-3 pages): itemize
  • Four-quadrant boundary summary
  • Rationale for each major scope decision (why in-scope? why out-of-scope?)
  • Expansion criteria with triggers
  • Stakeholder approvals itemize
  • Circulate to workshop participants for validation
  • Incorporate into Problem Definition Brief (A1.3 output)

Quality Criteria

Assess boundary definition quality using these criteria:

1. Completeness: All four quadrants defined with in-scope, out-of-scope, and expansion criteria

2. Specificity: Boundaries concrete, not abstract

  • ✗ Poor: "Focus on high-priority users"
  • ✓ Good: "Mid-market B2B sales managers (7-25 reps), exclude enterprise VPs"

3. Root cause alignment: In-scope functions/causes directly address validated root causes

  • Check: Does proposed scope address 2-4 high-leverage causes identified in Step 3?

4. Feasibility: In-scope elements within team capability

  • Check: Technical lead confirms in-scope causes addressable with available resources/timeline

5. Strategic coherence: Boundaries form coherent whole, not arbitrary collection

  • Check: Can explain to stakeholder "why this scope?" with clear rationale

6. Expansion criteria specificity: Clear triggers for future expansion, not vague "maybe later"

  • ✗ Poor: "Consider enterprise users in future"
  • ✓ Good: "Add enterprise VPs when SMB achieves 75% 6-month retention"

7. Stakeholder alignment: Key stakeholders explicitly agree (documented approval)

Complete Procedure

Step 1: State the Problem (Validated Need)

Begin with validated need statement from A1.3 Step 1, framed as observable problem:

Format: "[User type] experience [problem/obstacle] when [context], preventing [goal]"

Example: "Mid-market B2B sales managers struggle to confidently assess which deals will close this quarter when preparing weekly forecasts, leading to anxiety and defensive sandbagging"

Write problem statement at top of worksheet. This anchors analysis—ensures "why" questions trace back to validated user experience.

Step 2: Ask "Why" and Document Answer with Evidence

Ask: "Why does this problem occur?"

Answer must be:

  • Evidence-based: Traced to A1.2 data (interview quotes, observations, artifacts)
  • Specific: Not vague abstraction
  • Actionable level: Not "because users are bad at X" (blames user), instead "because current tools don't support X" (identifies addressable cause)

Document answer on worksheet with evidence citation:

|p5cm| Why LevelAnswerEvidence
Problem:Sales managers struggle to assess deal close likelihoodValidated in A1.2 (12/15 managers reported anxiety)
Why 1?They lack confidence in deal health assessments"I don't trust my gut anymore" (P03, P07, P11)

Step 3: Ask "Why" Again (Iterative)

For the answer just documented, ask "Why is that the case?"

Continue chain:

|p5cm| Why LevelAnswerEvidence
Why 1:They lack confidence in deal health assessments"I don't trust my gut anymore" (P03, P07, P11)
Why 2?Current assessments based on gut-feel, not validated by dataObserved: 8/8 managers use "feeling" language; no quantitative health metrics in CRM

Key principle: Each "why" goes deeper, from symptom toward structural/systemic cause. Don't lateral hop (different causes at same level)—go vertical (deeper cause chain).

Step 4: Repeat Until Root Cause Reached

Continue asking "why" until reaching root cause. Indicators you've reached root cause:

1. Actionability: Cause is something team/organization can address (not immutable constraint)

2. Structural/systemic: Cause is about systems, processes, tools, capabilities—not individual user failure

3. Solution-generative: Addressing this cause suggests clear solution direction

4. Explanation: This cause explains higher-level symptoms (causal logic sound)

5. No deeper "why": Asking "why" again leads to vague/unhelpful answers or dead-end

Example chain reaching root cause:

|p4.5cm| Why LevelAnswerEvidence
Problem:Managers struggle to assess deal close likelihoodA1.2: 12/15 managers reported
Why 1:Lack confidence in assessments"Don't trust gut" (P03, P07, P11)
Why 2:Assessments based on gut-feel, not data-validatedObserved: no quantitative health metrics
Why 3:CRM lacks structured deal health indicatorsCRM screenshots show stage/amount only
Why 4:Organization hasn't defined what "deal health" meansInterview P05 (sales ops): "No standard definition of health signals"
Why 5?[Stops here—"why hasn't organization defined?" leads to organizational history/culture, less actionable]

Root cause identified: "Organization hasn't defined what 'deal health' means" → Solution-generative (can define health framework), Structural (organizational capability gap), Actionable (within innovation project scope).

Stopping Criteria: When to Stop Asking "Why"

Stop when:

  • Reached actionable structural/systemic cause (criteria 1-5 above met)
  • Further "why" leads outside team/organization control (regulatory constraints, laws of physics, market forces)
  • Answer becomes vague/circular—"because that's how things are done" signals stopping point
  • Team consensus that current level is addressable and high-leverage

Don't stop if:

  • Still at proximate symptom level (user behavior description, not cause)
  • Answer blames user ("because users are lazy/stupid")—go deeper to understand structural enablers
  • Could ask more specific "why" (current answer too vague)
  • Haven't reached solution-generative level (no clear solution direction yet)

Note on "Five": The number "five" is not magical—Ohno found most problems reach root cause in 3-5 iterations. May reach root in 3 whys (simpler problem) or need 6-7 (complex). Focus on reaching actual root, not hitting exactly five.

Step 5: Validate Root Cause

Before concluding, validate that identified cause is genuine root:

Counterfactual test: If we addressed this root cause, would the problem (proximate symptom) go away or significantly diminish?

Example: If organization defines deal health framework → provides structured indicators → managers could base assessments on data → confidence increases → forecast anxiety reduces. ✓ Causal chain logical.

Alternative causes test: Are there other causes we haven't explored?

Run quick check:

  • "What else could cause this problem?"
  • If alternative causes exist, document as secondary causes (may need Ishikawa for multi-causation)
  • If single dominant cause, Five Whys sufficient

Stakeholder validation: Share why-chain with 2-3 users who experienced need. Do they agree with causal logic?

Script: "We found that [problem] occurs because [Why 1] because [Why 2] because [root cause]. Does this match your experience? Are we missing anything?"

Step 6: Document and Communicate

Create concise root cause summary:

  • Problem statement (validated need)
  • Root cause identified (final "why" answer)
  • Complete why-chain showing reasoning
  • Evidence supporting each level
  • Validation performed (counterfactual, alternatives checked, stakeholder feedback)

Communicate to stakeholders using simple format:

tcolorbox[title=Root Cause Summary Example,colback=alphacolor!5,colframe=alphacolor] Problem: Sales managers struggle to confidently assess deal close likelihood

Root Cause: Organization hasn't defined what "deal health" means—no standardized framework for assessing deal risk/strength

Why Chain:

  • Problem → Why? Managers lack confidence
  • Why lack confidence? → Assessments based on gut-feel
  • Why gut-feel? → CRM lacks health indicators
  • Why lacking? → Organization hasn't defined health framework

Evidence: Interview P05 (sales ops): "No standard definition"; CRM screenshots confirm no health metrics; 8/8 observed managers used "feeling" language

Implication: Solution should define deal health framework and surface indicators in workflow tcolorbox

Quality Criteria

1. Evidence-based: Each "why" answer supported by A1.2 data (quotes, observations), not speculation

2. Progressive depth: Each level goes deeper than previous—from symptom toward structural cause

  • ✗ Circular: "Why X? Because Y. Why Y? Because X" (no progression)
  • ✓ Progressive: "Why X? Because Y. Why Y? Because Z (deeper systemic issue)"

3. Avoids blame: Doesn't stop at "user is bad/lazy/stupid"—goes to structural enablers

  • ✗ Blame: "Why errors? Users don't pay attention" (dead-end)
  • ✓ Structural: "Why don't pay attention? Interface doesn't highlight critical fields" (actionable)

4. Solution-generative: Root cause suggests solution direction

  • ✗ Abstract: "Why? Company culture doesn't value quality" (vague, not actionable)
  • ✓ Specific: "Why? No quality checkpoints in workflow" (actionable—add checkpoints)

5. Validated: Counterfactual and stakeholder validation performed

6. Single-chain focus: Follows one dominant cause chain (doesn't branch—if multi-causation, use Ishikawa)

Complete Procedure

Step 1: Start with Validated Need (Middle of Ladder)

Place validated need statement in center of ladder—this is your starting point, neither top nor bottom.

Format: "[User type] struggle to [action] when [context] because [immediate obstacle]"

Example: "B2B sales managers struggle to assess deal close likelihood when preparing weekly forecasts because they lack confidence in their assessments"

Write this as "Level 0" in center of workspace. The ladder will extend both upward (why) and downward (how).

Step 2: Climb "Why" Ladder (Upward—More Abstract)

Starting from Level 0 (validated need), ask: "Why does this problem exist? What is the underlying cause?"

Document Level +1 (one level more abstract):

Example:

  • Level 0: Managers lack confidence in assessments
  • Level +1: Assessments based on gut-feel, not validated by data

Continue climbing. For each new level, ask "why" about the previous level:

  • Level +2: Why gut-feel? → CRM lacks structured deal health indicators
  • Level +3: Why lacking indicators? → Organization hasn't defined what "deal health" means
  • Level +4: Why undefined? → Company culture prioritizes speed over forecast rigor
  • Level +5: Why culture? → Market incentives reward revenue growth over operational excellence

When to stop climbing:

  • Reached causes outside organizational control (market forces, industry structure, regulations)
  • Cause becomes too abstract to act upon ("because capitalism" is not actionable)
  • Further "why" yields vague or circular answers
  • Typically 3-5 levels above starting point sufficient

Step 3: Descend "How" Ladder (Downward—More Concrete)

Return to Level 0 (validated need). Now ask: "How does this problem manifest? What are concrete observable symptoms?"

Document Level -1 (one level more concrete):

Example:

  • Level 0: Managers lack confidence in assessments
  • Level -1: Managers experience anxiety and second-guess forecasts

Continue descending. For each new level, ask "how" about the previous level:

  • Level -2: How does anxiety manifest? → Sunday evening stress, difficulty sleeping before Monday forecast calls
  • Level -3: How observed? → Managers report "forecast dread"; observed body language (P03 fidgeting with hands when discussing forecast)
  • Level -4: How affects behavior? → Defensive sandbagging—managers systematically underestimate to avoid missing forecast

When to stop descending:

  • Reached specific behavioral observations directly from A1.2 data
  • Further "how" becomes trivial detail (exact words said, specific timestamps)
  • Connected back to concrete user quotes and observed behaviors
  • Typically 2-4 levels below starting point sufficient

Visual Ladder Structure

At completion, ladder shows hierarchy:

verbatim Level +5: Market incentives (growth over operations) ↓ WHY Level +4: Company culture (speed over rigor) ↓ WHY Level +3: Org hasn't defined "deal health" ↓ WHY Level +2: CRM lacks structured health indicators ↓ WHY Level +1: Assessments based on gut-feel ↓ WHY —————————————————- Level 0: MANAGERS LACK CONFIDENCE IN ASSESSMENTS ← Validated need —————————————————- ↓ HOW Level -1: Managers anxious, second-guess forecasts ↓ HOW Level -2: Sunday evening stress, sleep issues ↓ HOW Level -3: Fidgeting, nervous body language observed ↓ HOW Level -4: Defensive sandbagging behavior verbatim

Step 4: Test Alternative Paths (Prevent Confirmation Bias)

Don't assume first path up ladder is only or best explanation. Test alternatives:

At Level +1, ask: "What else could cause Level 0?"

Example:

  • Alternative path: Managers lack confidence BECAUSE they're inexperienced (training gap)
  • Alternative path: Managers lack confidence BECAUSE deals have become more complex (market shift)

For each alternative, climb "why" ladder and evaluate:

  • Which path has stronger evidence from A1.2?
  • Which leads to more actionable root cause?
  • Are multiple paths valid? (If yes, multi-causation—may need Ishikawa)

If alternative path more compelling, follow that ladder. If original path strongest, continue with it but document alternatives considered (transparency).

Step 5: Select Actionable Root Cause Level

Ladder shows 8-12 levels (±4-5 from center). Not all levels equally suitable as "root cause" for solution design. Select level using criteria:

Criteria for actionable root cause level:

1. Solution-generative: Does this level suggest clear solution direction?

  • ✓ Level +3 (Org hasn't defined "deal health"): Solution = define health framework
  • ✗ Level +5 (Market incentives): Solution unclear, too abstract

2. Within capability: Can team/organization address this cause?

  • ✓ Level +2 (CRM lacks indicators): Within product team scope
  • ✗ Level +4 (Company culture): Outside product team control

3. High leverage: Will addressing this cause significantly alleviate need?

  • ✓ Level +3: If defined health framework → can build indicators → confidence improves (high leverage)
  • ✗ Level -1 (Anxiety): Treating symptom, not cause (low leverage)

4. Evidence-supported: Is this level well-supported by A1.2 data?

  • Check evidence citations for this level—solid data or speculation?

5. Stakeholder-validated: Do users/stakeholders agree this is core issue?

  • Share level with users: "We think core issue is X—does that resonate?"

Typical selection: Root cause usually 2-4 levels above validated need—deep enough to be structural/systemic, not so high as to be unactionable.

Example: Select Level +3 ("Organization hasn't defined what 'deal health' means") as root cause for solution design. Levels +4, +5 too abstract/outside control. Levels +1, +2 solvable but Level +3 more foundational (solving +3 enables solving +2, +1).

Step 6: Document Ladder and Root Cause Selection

Create visual ladder diagram (export from whiteboard/digital tool) and narrative explanation:

Ladder summary document should include:

  • Visual ladder showing all levels (why-path upward, how-path downward)
  • Evidence citations for each level (interview quotes, observations)
  • Alternative paths explored and why primary path selected
  • Root cause level identified (highlighted on ladder)
  • Selection rationale (5 criteria assessment)
  • Solution implications (how root cause informs solution direction)

Quality Criteria

1. Bidirectional completeness: Ladder climbs (why) AND descends (how)—both directions explored

2. Progression consistency: Each step changes abstraction level appropriately

  • Why-steps become more abstract (broader, more systemic)
  • How-steps become more concrete (specific behaviors, observations)
  • ✗ Lateral hop: Switching to different cause at same abstraction level

3. Evidence at all levels: Both abstract (why) and concrete (how) levels supported by data

  • Common error: high-level causes become speculation (no evidence)
  • Fix: Cite interviews/observations even for abstract levels

4. Alternative paths explored: Documented consideration of alternative causal explanations

5. Root cause selection justified: Clear explanation why specific level selected using 5 criteria

6. Validation performed: Counterfactual test and stakeholder validation completed

Complete Procedure

Preparation (Before Workshop)

1. Create fishbone diagram template

Visual structure:

verbatim Cause Category 1 / / ———/——— / / / / Cause Category 2/ / / / / [VALIDATED NEED] ←—● (Fishbone spine) Cause Category 3 —————— Cause Category 4 verbatim

Physical: Draw on poster (36"×48") or whiteboard with:

  • Horizontal spine (arrow pointing to problem box on right)
  • 4-6 diagonal "bones" (categories) branching from spine
  • Sub-bones for specific causes within each category

Digital: Create in Miro/Mural using fishbone template (most tools have built-in templates)

2. Select cause categories

Classic 6 M's (manufacturing context, Ishikawa's original):

  • Methods (processes, procedures)
  • Machines (equipment, technology)
  • Materials (inputs, resources)
  • Measurements (data, metrics, monitoring)
  • Manpower (people, skills, training)
  • Mother Nature (environment, external factors)

Customize categories for your context:

Software/SaaS context:

  • Process (workflows, procedures)
  • Technology (tools, systems, infrastructure)
  • Data (availability, quality, integration)
  • People (skills, incentives, capacity)
  • Organizational (structure, culture, policies)
  • External (market, customers, partners)

Healthcare context:

  • Clinical Process (care workflows, protocols)
  • Technology/EMR (health IT systems)
  • Regulatory/Compliance (policies, mandates)
  • Organizational/Administrative (staffing, budgets)
  • Cultural/Professional Norms (medical culture, hierarchy)
  • Patient/Environmental Factors (patient complexity, facility constraints)

Consumer product/service context:

  • User Behavior (habits, mental models, expectations)
  • Product Design (features, interface, affordances)
  • Context/Environment (usage situation, constraints)
  • Social Norms (peer influence, social acceptability)
  • Triggers/Reminders (cues for action)
  • Resources (time, money, cognitive load)

Category selection principles:

  • Mutually exclusive: Minimize overlap between categories (each cause fits primarily in one category)
  • Collectively exhaustive: Categories cover all potential cause domains
  • 4-6 categories optimal: Fewer than 4 too broad, more than 6 too fragmented
  • Context-appropriate: Match problem domain (don't use manufacturing categories for service design)

Label fishbone "bones" with selected categories before workshop.

3. Prepare evidence packets

For each participant, create:

  • Validated need statement (1 page)
  • A1.2 summary: key themes, user quotes, observation highlights (2-3 pages)
  • Interview transcript excerpts showing cause-related statements
  • Blank fishbone worksheet for note-taking

Distribute 2-3 days before workshop for pre-reading.

Workshop Execution (2-3 hours)

SECTION 1: Context Setting (15 minutes)

[0-5 min] Review validated need

Present need statement and evidence snapshot. Brief recap—assume participants read pre-work, don't re-present everything.

[5-10 min] Introduce Ishikawa method and categories

Explain:

  • Purpose: Map multiple contributing causes across categories
  • Structure: Each "bone" is a category; we'll brainstorm causes within each
  • Process: Diverge (generate many causes) → Converge (prioritize/validate)

Show fishbone structure with categories labeled.

[10-15 min] Ground rules

  • Evidence-based: Causes must trace to A1.2 data (cite interviews/observations)
  • Defer judgment: During brainstorm, generate causes—don't debate whether "right" (evaluate later)
  • Specificity: Avoid vague causes like "communication problems"—specify what aspect of communication
  • No blame: Frame structurally ("System lacks X") not personally ("People are bad at X")

SECTION 2: Cause Brainstorming by Category (90 minutes)

Approach: Work through categories sequentially (NOT all simultaneously—too chaotic)

For each category (15 minutes per):

[Minute 0-10] Silent brainstorm → Share

  • Participants individually write causes on sticky notes (one cause per note)—5 minutes silent
  • Round-robin sharing: each person places notes on fishbone under appropriate category, reads aloud—5 minutes

[Minute 10-12] Cluster similar causes

Group related sticky notes (duplicates, variations on same theme)

[Minute 12-15] Add sub-causes

For major causes, ask "What specific factors contribute to this?" Add sub-bones (causes of causes)

Example structure:

verbatim CATEGORY: Technology/CRM ├─ Main cause: CRM lacks deal health indicators │ ├─ Sub-cause: No risk scoring system │ ├─ Sub-cause: Stage doesn't reflect momentum │ └─ Sub-cause: Historical close rate not surfaced └─ Main cause: Data scattered across systems ├─ Sub-cause: Email signals not in CRM └─ Sub-cause: Customer data in separate platform verbatim

Repeat for all 4-6 categories (90 minutes total if 6 categories × 15 min each)

Facilitation tips:

  • Keep energy high—enforce timeboxes strictly
  • If category yields few causes (<3), consider merging with another or revisiting category selection
  • If category overflows (>15 causes), may need to split into subcategories

SECTION 3: Validation and Evidence Mapping (30 minutes)

[90-105 min] Evidence check

For each cause on diagram, facilitator asks: "What evidence supports this?"

Participant who added cause cites:

  • Interview quote (P03 said "...")
  • Observation (saw X behavior)
  • Artifact (CRM screenshot shows Y)

If no evidence → mark cause as "hypothesis—needs validation" (don't remove, but flag as speculative)

Quality threshold: ≥70% of causes should have direct A1.2 evidence. If <50%, insufficient discovery—may need additional A1.2 investigation.

[105-120 min] Cross-category patterns

Look for connections across categories:

  • Do certain user types (if segmented) associate with specific categories?
  • Do causes in one category amplify causes in another? (interactive effects)
  • Are there "root causes of root causes"—deeper factors appearing across multiple categories?

Mark these patterns on diagram (use dotted lines connecting causes, annotations)

SECTION 4: Cause Prioritization and Boundary Definition (45 minutes)

[120-145 min] Categorize causes by addressability

For each cause, assess whether addressable by innovation team:

3 buckets:

1. Addressable causes (in-scope):

  • Team can directly influence/solve
  • Within technical, organizational, budgetary capability
  • High leverage (addressing significantly alleviates need)
  • Mark with green dot or "A"

2. Constraint causes (accept as given):

  • Outside team control (regulatory, market forces, organizational culture/politics)
  • Technically/politically infeasible to change
  • Mark with red dot or "C"

3. Deferred causes (future/parallel):

  • Addressable in principle BUT requires capabilities not yet developed OR other team owns
  • Lower priority than addressable causes—tackle high-leverage first
  • Mark with yellow dot or "D"

[145-165 min] Prioritize addressable causes

Among addressable causes (green), prioritize using:

Impact: How much does this cause contribute to validated need? (High/Medium/Low)

Effort: How difficult to address? (Low/Medium/High effort)

Evidence strength: How well-supported by A1.2? (Strong/Moderate/Weak)

Create simple priority matrix or vote:

  • Each participant gets 5 votes (dot voting)
  • Place dots on highest-priority addressable causes
  • Top 3-5 causes become focus for solution design

Post-Workshop Documentation

Within 24-48 hours:

1. Export fishbone diagram

  • Photograph (if physical) or export (if digital)
  • Create clean digital version (Lucidchart, Visio, PowerPoint, Miro export)
  • Color-code causes: Addressable (green), Constraint (red), Deferred (yellow)

2. Create cause summary document (3-5 pages)

For each category:

  • List major causes with sub-causes
  • Evidence citations (interview/observation sources)
  • Addressability assessment (A/C/D)
  • Priority ranking for addressable causes

3. Define problem boundaries (feeds Step 4)

Ishikawa informs Boundary Definition Canvas (Method: Boundary Definition Canvas):

  • In-scope: Top 3-5 addressable causes
  • Out-of-scope: Constraint causes (accepted as given)
  • Deferred: Lower-priority addressable or parallel-team causes

This becomes Quadrant 4 (Root Cause Boundaries) of Boundary Canvas.

Quality Criteria

1. Category appropriateness: Categories match problem domain (not using generic 6 M's when domain-specific categories more appropriate)

2. Evidence-based: ≥70% of causes directly supported by A1.2 data (quoted, cited)

3. Comprehensiveness: Multiple causes per category (3-8 typical)—single cause per category suggests insufficient exploration

4. Specificity: Causes concrete, not vague

  • ✗ Poor: "Communication issues"
  • ✓ Good: "Sales and marketing don't share deal intelligence—marketing unaware of customer pain points discussed in sales calls"

5. Structural framing: Causes describe systemic/structural issues, not blame

  • ✗ Blame: "Managers are lazy about data entry"
  • ✓ Structural: "Data entry required during high-pressure moments (end of quarter), competing with urgent tasks"

6. Addressability assessment: Each cause categorized as A/C/D with clear rationale

7. Prioritization performed: Addressable causes ranked (don't just list—must prioritize)

8. Cross-category patterns identified: Connections across categories documented (not just siloed category analysis)

Digital vs. Physical Trade-offs

p5cmp5cm AspectPhysical (Sticky Notes)Digital (Miro/Mural)
Pros• Tactile engagement • Entire team works simultaneously • Spatial flexibility • Good for co-located teams• Remote collaboration possible • Persistent—easy to revisit/refine • Searchable • Easy to share/present • No space constraints
Cons• Difficult to preserve (photograph required) • Not accessible remotely • Space constraints • Manual documentation effort• Less tactile feel • Can be visually overwhelming • Requires tool familiarity • Simultaneous editing creates confusion
Best forCo-located teams, in-person synthesis workshopsDistributed teams, async collaboration, long-term reference

Common Challenges and Solutions

Challenge 1: Too many notes (>500), overwhelming

Solution: Two-phase clustering

  1. Phase 1: Each researcher clusters their own interviews (creates 50-80 notes per researcher, pre-organized)
  2. Phase 2: Team synthesizes researcher-level clusters (more manageable scale)

Challenge 2: Disagreement on clustering—notes fit multiple places

Solution:

  • Acknowledge that clustering has inherent ambiguity—no single "right" answer
  • Try both placements, see which creates more internally coherent cluster
  • If genuinely spans both, duplicate note (physically or digitally) in both clusters
  • Cross-reference: "See also [related cluster]"

Challenge 3: Dominant voices drive clustering (groupthink)

Solution:

  • Enforce silent clustering phase (Step 1)—prevents early influence
  • Facilitator actively solicits quieter participants: "We haven't heard from [name] on this cluster—what's your read?"
  • Vote on ambiguous placements if needed (majority rule with dissent noted)

Challenge 4: Analysis paralysis—endless refinement

Solution: Time-box each step

  • Step 1 (Silent clustering): 90 min maximum
  • Step 2 (Review): 45 min maximum
  • Step 3 (Labeling): 60 min maximum
  • Step 4 (Hierarchical): 90 min maximum
  • Step 5 (Patterns): 60 min maximum
  • Total: 5.5-6 hours—if extending, take breaks but enforce eventual closure

Perfect is enemy of good—synthesis should reveal major patterns, not achieve perfection.

Boundary Definition Canvas method:boundary-definition

Used in: • A1.3 Step 4 (Boundary definition) ~ Related Activities: • A3 (Concept scoping) • A6 (Validation scope)

Tools and Resources

Physical workshop:

  • Large poster (36"×24" minimum) with four quadrants pre-drawn
  • Sticky notes (3 colors: in-scope, out-of-scope, expansion criteria)
  • Markers for annotations
  • Camera for documentation

Digital workshop:

  • Miro, Mural, or FigJam board with boundary canvas template
  • Video conferencing for remote participants
  • Collaborative editing enabled for all participants

Template downloads:

  • Boundary Definition Canvas (PowerPoint, Miro template, printable PDF)
  • Workshop facilitation guide
  • Boundary documentation template

Sample Size / Duration

Participants: 5-8 people (decision-makers with authority to set scope)

  • Essential: Product owner, technical lead, business stakeholder
  • Optional: Design lead, operations representative, customer success
  • Avoid: >10 people (decision-making becomes unwieldy)

Duration:

  • Workshop: 90-120 minutes
  • Pre-work: 30 minutes (review need statement and root causes)
  • Post-workshop documentation: 2-3 hours
  • Total: 4-6 hours (primarily workshop facilitator time)

Common Challenges and Solutions

Challenge 1: Scope creep during workshop ("we should also include X")

Symptoms:

  • Every out-of-scope element gets "but what about...?" pushback
  • Scope expands to encompass too much
  • Workshop extends beyond 2 hours without closure

Solutions:

  • Return to need and root causes: "Does X directly address our validated need or high-leverage root cause? If not, it's adjacent priority."
  • Use expansion criteria: "X is valuable but not V1—let's document as expansion when [trigger]"
  • Enforce prioritization: "We can do A, B, C deeply OR A, B, C, D, E, F, G, H shallowly. Which creates more value?"
  • Timebox: If debate exceeds 10 minutes, table for follow-up bilateral discussion, continue workshop

Challenge 2: Disagreement between product and engineering on feasibility

Symptoms:

  • Product wants broad scope, engineering says infeasible
  • Technical complexity not understood by product stakeholders
  • Boundaries defined that engineering can't deliver

Solutions:

  • Pre-workshop technical feasibility assessment: Engineering reviews root causes, flags high-complexity items before workshop
  • Effort-impact matrix during workshop: For each candidate scope element, plot effort (engineering) vs. impact (product)—prioritize high-impact, low-effort
  • Phasing: If high-value but high-effort, agree on phased approach—simplified V1, full capability V2
  • Facilitator neutrality: Facilitator (ideally not product or engineering) balances competing concerns

Challenge 3: User/stakeholder boundaries too narrow, misses significant need prevalence

Symptoms:

  • Focus on single user type when need actually spans multiple types
  • Discovery (A1.2) included diverse users, but boundaries exclude most
  • Risk building for narrow niche that doesn't justify investment

Solutions:

  • Review A1.2 prevalence data: How many user types experienced need? What was intensity distribution?
  • Test universality: Can solution designed for Type A also serve Type B with minor adaptation? If yes, broaden boundaries.
  • Core-vs-edge framing: Design core for primary user type (80% of effort), accommodate secondary types at edge (20% of effort)

Challenge 4: Constraint causes incorrectly categorized as addressable

Symptoms:

  • Boundaries include causes outside team control (organizational culture, executive mandates, regulatory)
  • Scope assumes capabilities that don't exist
  • Over-optimism about what's changeable

Solutions:

  • Political feasibility check: For each cause, ask "Who has authority to change this? Will they?"
  • Conservative categorization: If doubt whether addressable, start in "constraint" or "deferred"—can always expand, harder to contract
  • Stakeholder validation: If cause involves other teams/departments, confirm with them before categorizing as addressable

Example: Sales Forecast Confidence

Context (from A1.2 and A1.3 Steps 1-3):

  • Validated need: B2B sales managers (7-25 reps) struggle to confidently assess which deals will close this quarter, leading to forecast anxiety and defensive sandbagging
  • Root causes identified: (1) Lack of structured deal health signals, (2) CRM shows stage not risk, (3) Gut-feel assessment not validated by data, (4) Sales comp incentivizes sandbagging

Completed Boundary Canvas:

|p5.5cm| Q1: User/Stakeholder BoundariesQ2: Functional Boundaries
In-scope: Mid-market B2B sales managers (7-25 reps), individual contributor sellers (deal owners)In-scope: Deal health assessment (core), deal risk identification, pipeline confidence scoring
Out-of-scope: Enterprise sales VPs (>100 reps—different workflow), Solo sellers (<5 reps—use lightweight tools), Sales ops/analysts (different job)Out-of-scope: Forecast aggregation/roll-up (separate tool), Executive reporting (different user), Retrospective win/loss analysis
Expansion: Add enterprise VPs when SMB achieves 75% weekly active usage + 70% forecast accuracy improvementExpansion: Add forecast roll-up when deal assessment achieves 80% manager adoption
Q3: Contextual BoundariesQ4: Root Cause Boundaries
In-scope: Weekly forecast prep (Sunday PM/Monday AM), Desktop/web environment, Individual analysis workflowAddressable (in-scope): (1) Lack of structured deal health signals—build signal framework, (2) CRM shows stage not risk—surface risk indicators
Out-of-scope: Real-time in-meeting forecast defense, Mobile/on-the-go (requires mobile app), Collaborative forecast sessions (team reviews)Constraints (accept): (4) Sales comp incentivizes sandbagging—outside product scope, compensation policy
Expansion: Add in-meeting support when individual analysis achieves 80% accuracy validationDeferred (future): (3) Gut-feel not validated—requires historical deal outcome data + ML, defer to V2
Expansion: Add ML validation when 12+ months deal outcome data collected

Key decisions and rationale:

  • Narrowed to mid-market: Enterprise VPs have different workflow (aggregate team forecasts, not individual deal assessment)—separate use case. Start focused.
  • Excluded forecast aggregation: Managers already have tools (Excel, CRM roll-ups) for aggregation—validated need is deal-level confidence, not aggregation
  • Desktop-only V1: Mobile adds substantial complexity (responsive design, offline sync), not core to validated need context (weekly prep at desk)
  • Addressed 2 of 4 causes: Signals and risk indicators directly solvable. Sandbagging is compensation issue (constraint). ML requires data infrastructure not yet available (defer).

Expansion strategy: Once V1 validates value for mid-market (adoption + forecast accuracy improvement), expand user boundaries to enterprise, add ML-powered validation, build mobile support for in-meeting use.

Data Processing Workflow method:data-processing

Used in:

  • A1.2 Step 2.6 (After data collection)
  • A1.3 (Root cause data)
  • A1.4 (Exploration data)
  • A6 (Validation feedback)
  • A7 (Pilot data)

Universal Application: All research-based activities

Complete Workflow (6 Steps)

Step 1: Recording and Initial Documentation

During interviews:

  • Audio record with explicit permission
  • Take brief notes (key quotes, observations, topics to probe) but don't try to transcribe—focus on conversation
  • Note non-verbal cues in brackets: [long pause], [frustrated tone], [laughs]

Immediately after interview (within 24 hours):

  • Write 1-page debrief memo: itemize
  • Key insights from this interview
  • Notable quotes
  • Follow-up questions that emerged
  • Initial hypotheses about user type itemize
  • Tag recording with participant code (P01, P02...) and date
  • Upload recording to secure storage

Step 2: Transcription

Transcription options:

p4cmp3cmp2.5cm ServiceProsConsCost
Otter.aiReal-time transcription, affordable, searchableAccuracy 80-85%, needs cleanup$10/mo (600 min)
Rev.comHigh accuracy (99%), human transcriptionMore expensive, 24hr turnaround$1.50/min
Zoom autoIncluded with Zoom, instantAccuracy 75-80%, heavy cleanupFree (if using Zoom)
DescriptTranscription + audio editingLearning curve$15/mo

Recommendation:

  • Budget constrained (<$500): Otter.ai for all interviews
  • Quality priority: Rev.com for most impactful 8-10 interviews, Otter for remainder
  • Speed priority: Otter.ai real-time during interview

Transcript cleanup (budget 30-45 min per 60-min interview):

  • Fix obvious errors (technical terms, proper nouns, domain language)
  • Add speaker labels consistently (Interviewer: / Participant:)
  • Insert non-verbal cues from notes: [pause], [laughs], [frustrated tone]
  • Don't need perfect punctuation—just readable
  • Preserve participant's actual language including filler words ("um," "like")—reveals thinking patterns

Step 3: Anonymization

Required before sharing transcripts with broader team or storing long-term:

  • Participant names → codes (P01, P02) or pseudonyms
  • Company names → [Company A], [Their organization], [Competitor X]
  • Specific products → [Tool they use], [Current system]
  • Colleagues mentioned → [Their manager], [Team member], [Sales VP]
  • Identifying details → [Major client], [Recent acquisition]

Create participant metadata file (separate from transcripts): verbatim P01: Director of Sales, mid-market SaaS, 8yrs experience, Team of 12 P02: Sales Manager, enterprise software, 15yrs, Team of 25 verbatim

Step 4: Organization and Metadata

Create structured file organization:

verbatim /A1.2-Investigation-[Signal-Name]/ /00-Planning/ investigation-plan.docx /01-Recordings/ P01-interview-2026-01-15.m4a P02-interview-2026-01-17.m4a ... /02-Transcripts/ P01-transcript-cleaned.docx P02-transcript-cleaned.docx ... /03-Debriefs/ P01-debrief-memo.docx ... /04-Observations/ Obs01-field-notes-2026-01-20.docx ... /05-Synthesis/ affinity-diagram-export.pdf themes-documentation.docx /06-Deliverables/ structured-opportunity-description.docx verbatim

Maintain participant tracking spreadsheet:

CodeDateMethodRoleUser TypeTranscript?
P012026-01-15InterviewSales DirType AYes
P022026-01-17InterviewSales MgrType BYes
P032026-01-20ObservationSales MgrType AField notes
...

Step 5: Initial Coding (Atomic Insight Extraction)

Systematically extract insights from transcripts for affinity diagramming.

Approach 1: Manual (spreadsheet method, <15 interviews)

Create spreadsheet:

p3cmp2cmp2cmp4cm SourceQuote/ObservationThemeUser TypeNotes
P01"I spend Sunday nights stressing about Monday's forecast call"Forecast anxietyType AHigh emotional intensity
P01"CRM shows stage but not deal health"Data gapAll typesSystem limitation
P02"Manually track signals in spreadsheet"Workaround burdenType ATime-consuming
...

Extract 10-20 insights per interview → 200-400 total for 15-20 interviews.

Approach 2: Qualitative Analysis Software (>15 interviews)

Software options: NVivo, Dedoose, Atlas.ti, MAXQDA

Basic workflow:

  1. Import all transcripts into software
  2. Create code categories (deductive: based on need hypothesis; or inductive: emergent from data)
  3. Code line-by-line or paragraph-level: itemize
  4. Highlight text segment
  5. Apply code(s): "Need statement," "Obstacle," "Workaround," "Emotion," "Context"
  6. Add memo with interpretation itemize
  7. Run query to extract all segments with specific code
  8. Export coded segments as atomic insights for affinity diagramming

Software comparison:

p3.5cmp3.5cmp2.5cm ToolBest ForLimitationsCost
NVivoLarge datasets, complex coding, academic rigorSteep learning curve, Windows/Mac only$400-1500/license
DedooseWeb-based, team collaboration, mixed methodsLess powerful than NVivo$15/user/mo
Atlas.tiVisual network analysis, theory buildingLearning curve$300-600/license
MAXQDAUser-friendly, multimedia codingLimited export options$350-700/license

Decision rule: Use software if >15 interviews OR complex multi-method investigation (interviews + observations + diary studies) OR organizational investment in building qualitative research capability.

Step 6: Quality Review and Bias Mitigation

Before synthesis workshop, conduct quality review:

Traceability check:

  • Can each extracted insight be traced back to source (interview, observation)?
  • Are participant codes and metadata accurate?
  • Are verbatim quotes accurately transcribed (not paraphrased)?

Triangulation verification:

  • Do interview findings align with observation findings?
  • Discrepancies (users say X but observations show Y) indicate either: itemize
  • Social desirability bias in interviews
  • Unrepresentative observation contexts
  • Gap between perceived need and actual behavior (actual is more reliable) itemize
  • Flag discrepancies for synthesis discussion

Bias mitigation strategies:

Confirmation bias (seeking evidence supporting hypothesis):

  • Extract both confirming AND disconfirming evidence
  • Explicitly document: "5 users mentioned [need], but 3 users said [need] wasn't significant for them"
  • Don't cherry-pick quotes supporting desired conclusion

Recency bias (overweighting last few interviews):

  • Process all interviews before synthesis (don't synthesize incrementally)
  • Review debrief memos chronologically—patterns early interviews identified still valid?

Vividness bias (overweighting emotionally compelling stories):

  • Balance emotionally salient anecdotes with frequency data
  • "One user had intense experience with X" ≠ "X is widespread need"
  • Weight both intensity AND prevalence

Researcher preference bias (favoring needs aligned with researcher interests):

  • Team-based synthesis reduces individual bias
  • Have researchers code each other's interviews (cross-validation)
  • Explicitly discuss "Are we seeing this because it's there, or because we want to see it?"

Timeline and Resources

Typical timeline for 15 interviews + 5 observations:

ActivityDurationCan Parallelize?
Transcription (if outsourced)2-3 daysYes (simultaneous)
Transcript cleanup8-10 hoursYes (split across team)
Anonymization3-4 hoursYes
Initial coding / insight extraction12-15 hoursYes (each researcher codes own)
Quality review2-3 hoursNo (requires completeness)
Total elapsed:5-7 days

Resources:

  • Transcription: $300-1,800 (depending on service)
  • Software (if used): $15-50/month
  • Personnel: 25-35 hours researcher time (can distribute across 2-3 people)

Diary Studies method:diary-studies

Used in:

  • A1.2 Step 2.4 (Need validation)
  • A7 Step 3 (Usage patterns)

Best for: Distributed, infrequent, or private context needs

Study Design

Duration

  • 1 week: Sufficient for frequent needs (daily/multiple daily occurrences)
  • 2 weeks: Better for less frequent needs (few times per week), allows pattern stabilization
  • 3-4 weeks: For very infrequent needs, BUT participation fatigue becomes serious concern—diminishing compliance

Recommendation: Default to 2 weeks unless strong reason for shorter/longer. Longer studies require more participant management effort.

Logging Frequency

Three approaches:

1. Event-based logging: "Log whenever need-related experience occurs"

  • Pros: Most accurate, captures experiences in moment
  • Cons: Requires participant vigilance, irregular timing
  • Best for: Infrequent but salient experiences (user will remember to log)

2. Time-based logging: "Log at set intervals (end of day, end of week)"

  • Pros: Easier compliance (predictable schedule), less burdensome
  • Cons: Memory decay—may miss details or forget experiences entirely
  • Best for: Frequent experiences (can recall multiple from day), less time-sensitive contexts

3. Hybrid: "Event-based for salient experiences + daily summary"

  • Pros: Balances accuracy and compliance
  • Cons: Dual burden (both event and daily)
  • Best for: Mixed frequency needs—some experiences very salient (log immediately), others less so (daily recap)

Recommendation: Event-based preferred for accuracy, but requires motivated participants and salient need. Use time-based if compliance concerns or less emotionally-charged experiences.

Logging Mechanism

Option 1: Mobile app (ExpiWell, Indeemo, custom app)

  • Pros: Push notifications (reminders), media capture (photos/video), timestamps automatic, data aggregates easily
  • Cons: Requires app download/setup, technical barriers for some users, cost
  • Cost: $500-2,000 for 2-week study with 10 participants

Option 2: Online form (Google Forms, Typeform, Qualtrics)

  • Pros: No app required, accessible from any device, free or low-cost, easy setup
  • Cons: No push notifications (must email reminders), less rich media, manual timestamp
  • Cost: Free (Google Forms) to $50/month (Typeform Pro)

Option 3: Messaging platform (WhatsApp, Telegram group)

  • Pros: Familiar interface, easy media sharing, conversational feel enables researcher probing, social accountability (group sees others posting)
  • Cons: Less structured data, privacy concerns (group visibility), data extraction manual
  • Cost: Free

Option 4: Physical journal

  • Pros: No technical barriers, appropriate for contexts where digital inappropriate (sensitive medical situations, elderly users)
  • Cons: Cannot enforce compliance (no reminders), delayed data access (must collect at end), manual transcription required
  • Cost: $5-10/participant for journal

Selection guide:

  • Tech-savvy users, budget available: Mobile app
  • Broad user base, budget-conscious: Online form + email reminders
  • Small sample (5-8), need rich interaction: WhatsApp group
  • Low-tech users or sensitive contexts: Physical journal

Diary Entry Prompts

Need-experience entry template:

tcolorbox[title=Diary Entry Prompts,colback=alphacolor!5,colframe=alphacolor] When? [Date/time auto-captured if digital, or "Morning/Afternoon/Evening"]

What were you trying to do? [Goal/job]

What happened? [Brief description of experience]

What got in your way? [Obstacle/need]

How did you handle it? [Coping strategy/workaround]

How did you feel? [Emotion: frustrated, anxious, satisfied, relieved, etc.]

Photo or screenshot (optional): [Visual context if relevant] tcolorbox

Design principles:

  • Keep prompts concise: Target 2-3 minutes per entry. Lengthy requirements cause abandonment.
  • Use simple language: Avoid jargon ("What were you trying to do?" not "What was your goal-oriented task objective?")
  • Make most fields required: Ensures data completeness, but have 1-2 optional (photo) for flexibility
  • Include emotion prompt explicitly: Users often skip emotional dimension unless explicitly asked

Daily summary prompts (if using time-based or hybrid):

tcolorbox[title=End-of-Day Summary,colback=alphacolor!5,colframe=alphacolor] How many times today did you experience [need-related challenge]?

Which experience was most challenging/frustrating today? [Signals intensity]

Anything else relevant we should know? tcolorbox

Participant Management

Recruitment

Diary studies require higher commitment than single-session interviews (1-2 weeks vs. 1 hour). Offer appropriate compensation to reflect burden:

  • Consumer participants: 100-200 for 1-week study, 150-300 for 2-week
  • B2B professionals: $250-400 for 2-week study
  • Hard-to-reach specialists: $400-600

Screen for commitment during recruitment:

  • "This study requires logging experiences daily for 2 weeks—about 5-10 minutes per day. Can you commit to that?"
  • Higher compensation attracts more committed participants

Onboarding

Onboarding call or video (15-20 minutes):

  1. Explain purpose: "We're trying to understand when and how [need] comes up in your daily life. Your entries help us see patterns."
  2. Walk through logging mechanism: itemize
  3. If app: Have participant download during call, complete practice entry together
  4. If online form: Send link, have them complete practice entry, confirm receipt
  5. If physical journal: Review template structure, show example filled entry itemize
  6. Set expectations: itemize
  7. Frequency: "We're hoping for at least [one entry every other day / daily summary / logging whenever X happens]"
  8. Entry length: "Brief is fine—2-3 sentences per prompt, shouldn't take more than 3 minutes"
  9. Duration: "[1 or 2] weeks, ending on [date]" itemize
  10. Provide contact: "If you have questions or technical issues, email/text me: [contact]"
  11. Confirm understanding: "Any questions about what we're asking you to do?"

Ongoing Engagement

Daily reminders (essential for compliance):

  • Timing: Send at consistent time matching user's typical experience (morning for work-related, evening for home/life)
  • Message examples: itemize
  • "Morning! Remember to log any [need-related experiences] today 📝"
  • "End-of-day check-in: Did [challenge] come up today? Log it if so!"
  • "You're halfway through week 1—thanks for your entries! 🎉" itemize
  • Channel: Push notification (if app), SMS (if provided phone numbers), email (if form-based)

Mid-week check-in:

Day 3-4 of each week, send personalized message:

  • "How's it going with the diary study? Any questions?"
  • "Noticed you logged [specific experience]—that's really helpful, thanks!"
  • If user hasn't logged recently: "Haven't seen an entry in a few days—everything okay? Let me know if you're running into issues."

Acknowledge entries:

When participant logs particularly detailed or insightful entry, respond:

  • "Thanks for that entry about [topic]—really helpful detail!"
  • Shows you're paying attention, reinforces participation
  • Don't acknowledge every entry (too burdensome for researcher), but 1-2 per week per participant

Compliance Challenges

Expected compliance: 60-80

Common issues and responses:

Participants forget to log:

  • Increase reminder frequency (twice daily instead of once)
  • Reminder message: "Quick reminder—if [need situation] happened today, log it now before you forget the details!"

Participant fatigue (entries become cursory after Week 1):

  • Mid-study encouragement: "You're past the halfway point—your entries have been really valuable, keep it up!"
  • Reduce burden if needed: "If daily is too much, aim for every other day—better to have fewer detailed entries than many quick ones"

Life interruptions (participant gets sick, travels, busy week):

  • Flexible accommodation: "I see you haven't logged in several days—if something came up, no problem. Resume when you can."
  • Extend study duration by a few days if needed to compensate for gap

Dropouts:

  • Some participants will drop out despite best efforts—accept this
  • If dropout happens early (first 3 days), recruit replacement participant
  • If dropout happens late (after 7+ days), may still have usable partial data

Data Analysis

Aggregate entries across participants:

Frequency analysis:

  • How often does need arise? Calculate: Total entries reporting need / Total participant-days
  • Example: 45 entries across 10 participants over 14 days = 45/(10×14) = 32
  • Per-participant variation: Does need arise daily for some but rarely for others? (Signals user type differences)

Context patterns:

  • When/where does need most commonly manifest? itemize
  • Time of day: Morning/afternoon/evening? Day of week?
  • Location: Home/work/transit?
  • Trigger situations: Identify common precipitating events itemize
  • Create context frequency table showing where entries cluster

Intensity distribution:

  • What percentage of experiences are high-frustration vs. low-frustration?
  • Code entries as: High intensity (strong emotional language, major consequences), Medium, Low
  • If >50

Coping strategies:

  • What workarounds appear repeatedly?
  • Catalog all coping strategies mentioned, frequency of each
  • Elaborate workarounds (multi-step, time-consuming) indicate high need significance

Individual participant analysis:

User type differentiation:

  • How does need experience vary across participant characteristics?
  • Segment participants by role, experience level, context—how does need frequency/intensity differ?
  • Informs user segmentation (2x2 matrix) in synthesis

Temporal patterns per user:

  • Does need intensity escalate over time (cumulative burden) or remain stable?
  • Plot intensity ratings over 14 days per participant—look for trends

Rich examples for documentation:

Select particularly detailed, emotionally-salient entries for inclusion in Structured Opportunity Description—diary entries often provide vivid illustration of need manifestation.

Example rich entry: quote "Tuesday evening, 7pm. Tried to make dinner plan for week but couldn't remember what ingredients we already have. Stood in pantry for 10 minutes checking shelves, still forgot half the things. Felt really frustrated—this happens EVERY week and I waste so much time. Ended up making yet another list on my phone that I'll probably lose." [P06, Day 4] quote

Sample Size

Target: 5-10 participants for 1-2 weeks

Fewer participants than interviews because each diary study generates many data points:

  • 1 participant logging 2×/day for 2 weeks = 28 entries
  • 8 participants = 224 total entries (comparable to 15-20 interviews in data volume)

Distribution:

  • If user types identified: 2-3 participants per type minimum
  • Ensures representation across variation in need experience

Empathy Mapping method:empathy-mapping

Used in:A1.2 Step 4(Synthesis, optional)

[Complete 2,000-word content with quadrant definitions, creation process, examples...]

Governance Decision Brief method:governance-brief

Used in:A1.2 Step 5(Documentation)

[Complete 1,500-word content with template, examples...]

Five Whys Analysis method:five-whys

Used in: • A1.3 Step 3 (Root cause analysis) ~ Origin: Toyota Production System (Taiichi Ohno, 1950s)

Tools and Resources

Five Whys Worksheet Template:

Simple table format (whiteboard, Word doc, spreadsheet):

|p5cm| LevelWhy Question → AnswerEvidence / Source
Problem:[Validated need statement][A1.2 validation data]
Why 1?
Why 2?
Why 3?
Why 4?
Why 5?
Root Cause:[Identified root—most actionable level]

Software/digital tools:

  • Google Sheets/Excel: Simple table
  • Miro/Mural: Visual flow diagram
  • Whimsical: Flowchart format
  • Notion/Confluence: Documentation with evidence links

No specialized software required—method designed for simplicity.

Sample Size / Duration

Participants: 2-4 people ideal

  • Minimum: 1 person (can conduct solo with evidence review)
  • Optimal: 2-3 (diverse perspectives reduce bias, validate logic)
  • Maximum: 5 (larger groups become unwieldy for focused analysis)

Duration:

  • Analysis session: 30-45 minutes (simple problem) to 60-90 minutes (complex)
  • Evidence review (pre-work): 30 minutes (reviewing A1.2 transcripts/notes)
  • Documentation: 30 minutes (creating summary, visual)
  • Total: 1.5-3 hours

Common Challenges and Solutions

Challenge 1: Stopping too early (proximate causes)

Symptoms:

  • Reach answer after 2 whys and stop
  • Answer is user behavior description, not structural cause
  • Solution direction not clear

Example:

  • Problem: Forecasts inaccurate
  • Why? Managers overestimate deal close likelihood
  • [Stops here] ← TOO EARLY (describes behavior, not cause)

Solutions:

  • Ask "why" about behavior: "Why do managers overestimate?" (don't accept behavior as endpoint)
  • Check solution-generative criterion: Can we design solution from this answer? If unclear, go deeper
  • Structural test: Is this about individual failure or systemic/structural issue? If individual, go deeper to find structural enabler

Challenge 2: Climbing too high (unsolvable systemic causes)

Symptoms:

  • Reach organizational culture, market forces, regulatory constraints (outside control)
  • Answer too abstract to act on
  • Team feels stuck ("we can't change that")

Example:

  • Why 4: No deal health framework defined
  • Why 5: Company culture doesn't prioritize forecast rigor
  • Why 6: CEO focused on growth over operational excellence ← TOO HIGH (outside project scope)

Solutions:

  • Stop at highest actionable level: When next "why" leaves your control, previous level is root cause for your purposes
  • Constraint framing: If reached constraint (culture, policy), accept as given and solve within constraint—"Given company culture is X, how can we address need?"
  • Reframe: "We can't change culture, but we can define health framework for our team/division" (scope appropriately)

Challenge 3: Confusing correlation with causation

Symptoms:

  • "Why" answer describes co-occurring factor, not actual cause
  • Causal logic breaks under counterfactual test

Example:

  • Problem: Forecasts inaccurate
  • Why: Managers work remotely
  • [Assumes remote work causes inaccuracy—correlation, not necessarily causation]

Solutions:

  • Counterfactual test: If managers worked in office, would forecasts be accurate? (If no, remote work isn't root cause)
  • Mechanism test: How does X cause Y? Can we articulate causal mechanism? If unclear, may be correlation
  • Alternative causes: What else could explain Y? If multiple equally plausible causes, further investigation needed

Challenge 4: Skipping evidence (speculation)

Symptoms:

  • "Why" answers based on assumptions, not A1.2 data
  • Team debates what "probably" causes problem without evidence
  • Root cause reflects team's prior beliefs, not user reality

Solutions:

  • Evidence requirement: For each "why" answer, facilitator asks "Where's the evidence?" Must cite interview, observation, artifact
  • Flag assumptions: If no evidence, mark answer as "hypothesis—needs validation" and conduct follow-up research
  • Return to A1.2 data: If stuck, pause Five Whys, review transcripts/notes for relevant evidence, resume analysis

Challenge 5: Multi-causation ignored (forced single-chain)

Symptoms:

  • Team debates which cause to follow—multiple plausible chains
  • Single chain feels incomplete, missing important contributing factors
  • Different stakeholders advocate different "why" paths

Solutions:

  • Acknowledge multi-causation: "This problem has multiple contributing causes"
  • Document alternative chains: Run Five Whys for 2-3 different causal paths
  • Switch methods: Use Ishikawa diagrams (designed for multi-causation) instead of Five Whys

Example: Meditation App Habit Formation

Context:

  • Validated need (A1.2): Stressed professionals want to build consistent meditation habit but abandon after 1-2 weeks despite motivation
  • A1.2 evidence: 18 interviews, 6 diary studies (2-week tracking)

Five Whys Analysis:

|p5.5cm| LevelAnswerEvidence
Problem:Professionals abandon meditation practice after 1-2 weeksA1.2: 14/18 interviewees reported abandonment; diary studies show 5/6 participants dropped off by day 10
Why 1?They don't experience sufficient progress/benefit to sustain motivationInterview P03: "Wasn't sure it was working"; P07: "Didn't feel different"; Diary study: no progress markers noted
Why 2?Meditation benefits are gradual and internal—hard to perceive without explicit trackingInterview P12 (meditation teacher): "Benefits accumulate slowly, people expect immediate change"; Literature: meditation effects emerge over 8+ weeks
Why 3?Current app lacks progress indicators—no visibility into benefit accumulationApp screenshots: shows streak count only, no other progress metrics; Interview P05: "Just shows I did 7 days, but so what?"
Why 4?App design assumes motivation is sustained by activity alone, not by visible progressObserved: app focuses on content delivery (guided sessions), not behavior change scaffolding; No progress dashboard evident
Root Cause:App lacks explicit progress and benefit visualization—doesn't scaffold gradual behavior change with progress feedbackBehavioral science lit: visible progress critical for habit formation (Fogg Behavior Model, BJ Fogg); Apps treating meditation as content delivery vs. habit-building miss progress mechanism

Validation:

Counterfactual test: If app provided progress visibility (e.g., weekly reflection scores, physiological indicators like HRV, benefits timeline), would users persist longer?

Hypothesis: Yes—users would see accumulating evidence of benefit, sustaining motivation through early weeks when subjective experience is ambiguous.

Alternative causes considered:

  • Time constraints? (No—diary studies show users had time, issue was motivation loss)
  • Content quality? (No—users praised guided sessions, issue wasn't dislike of content)
  • Reminder/trigger issues? (Secondary—some users mentioned forgetting, but primary issue was "not seeing benefit")

Stakeholder validation: Shared why-chain with 3 users who abandoned. Responses:

  • P03: "Yes, I had no idea if it was working—that's exactly why I stopped"
  • P07: "If I'd seen some progress, even small, I'd have kept going"
  • P11: "The streak number didn't mean much without knowing whether it was helping"

Solution direction: Design progress dashboard showing:

  • Trend in self-reported stress/focus over time
  • Physiological indicators (HRV, sleep quality) if integrates with wearables
  • Weekly reflection prompts with longitudinal view
  • "Benefits timeline" showing typical emergence of meditation effects (managing expectations)

Reframed need statement (incorporating root cause):

Original: "Stressed professionals want to build consistent meditation habit but abandon after 1-2 weeks"

Root-cause informed: "Stressed professionals need visibility into meditation progress and benefit accumulation to sustain motivation through gradual behavior change process"

This reframing shifts focus from "deliver meditation content" to "scaffold habit formation with progress feedback"—fundamentally different solution direction.

How-Why Laddering method:how-why-laddering

Used in: • A1.3 Step 3 (Root cause analysis) ~ Best for: Hierarchical causation—exploring causes at multiple abstraction levels

Tools and Resources

Physical workspace:

  • Large whiteboard or wall space
  • Sticky notes for each ladder level (different colors for why/how levels)
  • Markers for connections and annotations
  • Camera for documentation

Digital workspace:

  • Miro, Mural, Whimsical—visual diagramming tools
  • Ladder template (vertical layout with bidirectional arrows)
  • Ability to attach evidence/notes to each level

Template structure:

Create template with:

  • Center box for Level 0 (validated need)
  • 5 boxes above for "why" levels (+1 to +5)
  • 4 boxes below for "how" levels (-1 to -4)
  • Arrows connecting levels with "WHY" and "HOW" labels
  • Evidence sidebar for documenting sources

Sample Size / Duration

Participants: 2-4 people (same as Five Whys)

  • Facilitator (guides ladder climbing/descending, probes for specificity)
  • 1-3 domain experts or researchers (provide evidence, validate logic)

Duration:

  • Analysis session: 60-90 minutes itemize
  • Climb why-ladder: 25-35 min
  • Descend how-ladder: 20-30 min
  • Test alternatives: 15-20 min
  • Select root cause level: 10-15 min itemize
  • Evidence review (pre-work): 30-45 minutes
  • Documentation: 45-60 minutes (creating visual, narrative)
  • Total: 3-4 hours

Common Challenges and Solutions

Challenge 1: Confusing lateral exploration with vertical abstraction

Symptoms:

  • "Why" steps list multiple causes at same level rather than going deeper
  • Ladder branches sideways instead of climbing
  • Example: Level +1 lists "A, B, C, D" causes (lateral) instead of asking why about single path (vertical)

Solutions:

  • Focus on dominant path: Select strongest/most-evidenced cause and follow that path upward—explore alternatives in Step 4, not while climbing
  • Abstraction check: Ask "Is this more fundamental/systemic than previous level?" If no, try again
  • Facilitator intervention: "We're listing alternatives at same level—let's pick primary path and go deeper"

Challenge 2: Skipping "how" ladder (only climbing "why")

Symptoms:

  • Team climbs why-ladder thoroughly but skips descending
  • High-level causes disconnected from observed user experience
  • Risk of abstract theorizing not grounded in evidence

Solutions:

  • Mandate "how" descent: Facilitator enforces—"We've climbed, now we must descend"
  • Validation purpose: Explain that descending validates high-level causes actually manifest in user behavior observed in A1.2
  • Evidence grounding: "How" ladder should reconnect to specific A1.2 quotes/observations—if can't, high-level cause may be speculative

Challenge 3: Selecting root cause level based on preference, not criteria

Symptoms:

  • Team selects cause they "want" to address (interesting, aligns with existing roadmap)
  • Skips systematic assessment against 5 criteria
  • Selected level not actually best leverage point

Solutions:

  • Criteria scorecard: Create table scoring each level against 5 criteria (solution-generative, capability, leverage, evidence, validated)
  • Facilitator neutrality: Facilitator has no stake in which level selected—pushes for objective assessment
  • Stakeholder validation: Share multiple candidate levels with users—which resonates most?

Challenge 4: Evidence thins at higher abstraction levels

Symptoms:

  • Lower levels (concrete) have rich evidence, upper levels (abstract) based on assumptions
  • "Why" answers at +4, +5 are team's theories, not user-validated
  • Risk of confirmation bias (fitting observations to preferred theory)

Solutions:

  • Flag speculation: Mark levels with weak evidence as "hypothesis—validation needed"
  • Inference explicit: Distinguish direct evidence ("users said X") from inference ("X suggests underlying Y")
  • Validation follow-up: If selected root at high level with weak evidence, conduct follow-up interviews to validate before committing to solution

Example: Sales Forecast Confidence (Full Ladder)

Context: Mid-market B2B sales managers, forecast confidence need validated in A1.2

Complete Ladder:

|p4.5cm| LevelCause/ManifestationEvidence
3|c|WHY PATH (CLIMBING)
+5Market/industry incentivizes growth velocity over operational rigorIndustry analysis: SaaS valuation multiples favor growth rate; compensation tied to revenue targets
+4Company culture prioritizes speed-to-market over forecast accuracyInterview P14 (sales ops): "Leadership wants aggressive targets, not precision"
+3yellow!25Organization lacks standardized definition of "deal health" signalsyellow!25Interview P05 (sales ops): "No standard"; P08: "Everyone uses different criteria"
+2CRM lacks structured deal health indicatorsCRM screenshots: stage and amount only; no risk/health metrics
+1Assessments based on gut-feel, not data-validatedObserved: 8/8 managers used "feeling" language; Interview P03: "Go with my gut"
3|c|VALIDATED NEED (STARTING POINT)
0Managers struggle to confidently assess which deals will closeA1.2: 12/15 managers reported anxiety; all described forecast as "stressful"
3|c|HOW PATH (DESCENDING)
-1Managers experience anxiety and second-guess their forecastsInterview P03: "Constantly questioning myself"; P07: "Never sure if I'm right"
-2Sunday evening stress, difficulty sleeping before Monday forecast callInterview P11: "Dread Sunday nights"; P06: "Can't sleep, running scenarios in my head"
-3Observed body language: fidgeting, nervous energy when discussing forecastObservation: P03 fidgeting with pen during forecast discussion; P09 avoiding eye contact
-4Defensive sandbagging—systematic underestimation to avoid forecast missInterview P12: "Better to beat lowball than miss aggressive"; Observed: 6/8 managers sandbagged 10-20%

Alternative Paths Explored:

At Level +1, considered alternative:

  • Alt: Managers lack confidence BECAUSE inexperienced (training gap)
  • Evidence check: 11/15 managers had >5 years experience—not training issue
  • Verdict: Primary path (gut-feel/lack of data) stronger

Root Cause Selection:

Evaluated levels +2, +3, +4 against criteria:

LevelSolution?Capability?Leverage?Evidence?Valid?Score
+4 (Culture)✗ Vague✗ Outside scope✓ High⚠ Inference? Not tested2/5
yellow!25+3 (No health definition)yellow!25✓ Define frameworkyellow!25✓ Product scopeyellow!25✓ Highyellow!25✓ Strongyellow!25✓ Validatedyellow!255/5
+2 (CRM lacks)✓ Add indicators✓ Within scope⚠ Medium✓ Strong✓ Validated4/5

Selected: Level +3 ("Organization lacks standardized definition of 'deal health'")

Rationale:

  • Solution-generative: Define deal health framework (signals, thresholds, scoring)
  • Within capability: Product + sales ops collaboration can define framework
  • High leverage: Defining health enables indicators (+2), which enables data-based assessment (+1), which addresses confidence (0)—solves multiple levels
  • Strong evidence: Sales ops confirmed, multiple managers referenced inconsistency
  • Validated: Shared with 3 managers—all agreed "that's the core issue"

Solution Implication:

Root cause informs solution design focus:

  • Primary: Define standardized deal health framework (scoring rubric, risk signals, strength indicators)
  • Secondary: Surface framework in CRM workflow (indicators visible during deal assessment)
  • Tertiary: Track accuracy over time (validate that framework improves confidence)

If had selected +2 (CRM lacks indicators) instead: Would jump to building interface without defining what to show—cart before horse.

If had selected +4 (culture): Would attempt culture change program (outside scope, low leverage for product team).

Laddering enabled comparing levels explicitly and selecting strategically.

Investigation Planning method:investigation-planning

Used in:A1.2 Step 1(Planning)

[Complete 2,500-word content with templates...]

Ishikawa (Fishbone) Diagrams method:ishikawa-diagrams

Used in: • A1.3 Step 3 (Root cause analysis) ~ Also known as: Fishbone diagram, Cause-and-Effect diagram ~ Origin: Kaoru Ishikawa (1960s), quality management

Tools and Resources

Physical workshop:

  • Large poster (36"×48") or whiteboard
  • Sticky notes (3"×3") in single color for causes
  • Markers for category labels and sub-bones
  • Dot stickers (green/red/yellow) for addressability marking
  • Camera for documentation

Digital workshop:

  • Miro, Mural, Lucidchart (fishbone templates available)
  • Video conferencing for remote facilitation
  • Collaborative editing for all participants

Software for clean diagram export:

  • Lucidchart (professional fishbone templates)
  • Microsoft Visio (cause-effect diagram templates)
  • PowerPoint/Keynote (SmartArt fishbone)
  • Canva (visual design if presenting to executives)

Sample Size / Duration

Participants: 6-10 people

  • Cross-functional team (product, engineering, design, operations, business)
  • Include A1.2 researchers (essential—direct evidence access)
  • Include domain experts (understand systemic factors)
  • Avoid: >12 people (too large for productive brainstorming)

Duration:

  • Pre-work: 30-45 minutes (reviewing evidence packet)
  • Workshop: 2.5-3 hours itemize
  • Context: 15 min
  • Brainstorming (6 categories × 15 min): 90 min
  • Validation: 30 min
  • Prioritization: 45 min itemize
  • Post-workshop documentation: 3-4 hours (clean diagram, summary document)
  • Total: 6-8 hours (workshop facilitator/documenter)

Common Challenges and Solutions

Challenge 1: Category overlap—causes fit multiple categories

Symptoms:

  • Debate during workshop: "This belongs in Technology" vs. "No, it's Process"
  • Same cause appears in multiple categories (duplication)
  • Category boundaries unclear

Solutions:

  • Primary placement rule: Place cause in category where it's most directly addressed—if spans multiple, choose primary
  • Cross-reference: Add annotation: "See also [other category]" rather than duplicating
  • Revise categories: If systematic overlap (many causes span 2+ categories), categories may need refinement—pause workshop, revise category structure

Challenge 2: Category imbalance—some categories have many causes, others few

Symptoms:

  • Technology category has 15 causes; People category has 2
  • Suggests either: (a) real imbalance (problem truly concentrated in one domain), or (b) team bias (overemphasizing familiar domain)

Solutions:

  • Probe sparse categories: Facilitator asks "Are we missing causes in [sparse category]? Let's revisit evidence..."
  • Diverse team: Ensure cross-functional representation—engineers spot technical causes, ops spot process causes, designers spot UX causes
  • Accept real imbalance: If after probing, imbalance persists and evidence supports it, accept that problem concentrated in specific domain

Challenge 3: Abstraction mismatch—causes at different levels

Symptoms:

  • Some causes very specific ("Field X mislabeled in CRM"), others abstract ("Poor data governance")
  • Difficult to compare or prioritize causes at different abstraction levels

Solutions:

  • Use sub-bones: Abstract causes become main bones, specific causes become sub-bones underneath
  • "Why" within categories: For abstract cause, ask "What specifically causes this?"—creates hierarchy
  • Consistent abstraction target: Facilitator guides toward "actionable specificity"—specific enough to address, not trivial minutiae

Challenge 4: Confirmation bias—team sees causes they expect

Symptoms:

  • Ishikawa confirms team's pre-existing theory about problem
  • Causes that don't fit narrative dismissed or under-represented
  • Lack of surprising/unexpected causes (suggests insufficient exploration)

Solutions:

  • Return to evidence: If causes suspiciously aligned with team preferences, re-review A1.2 transcripts—what did users actually say?
  • Devil's advocate: Assign participant to argue for alternative causal explanations
  • External validation: Share diagram with users/stakeholders outside team—do they see missing causes?

Challenge 5: Constraint causes dominate—everything seems unaddressable

Symptoms:

  • Most causes marked red (constraint)—team feels helpless
  • Organizational/cultural factors dominate, technical/product factors minimal
  • Risk of analysis paralysis ("we can't solve anything")

Solutions:

  • Reframe constraints as boundaries: "We can't change X, but we can design solution that works within X"—constraint becomes design parameter
  • Deeper exploration: If all causes seem constraining, likely at too-high abstraction level—go more specific to find addressable sub-causes
  • Adjacent addressability: Even if can't address cause directly, can address its effects—mitigate rather than eliminate

Example: Healthcare Clinical Documentation Burden

Context:

  • Validated need: Primary care physicians spend 35-45% of patient visit time on EMR documentation during encounter, reducing eye contact and perceived empathy
  • A1.2: 18 physician interviews, 12 observation sessions (shadowing in clinic)

Selected Categories (healthcare context):

  1. Clinical Process
  2. Technology/EMR
  3. Regulatory/Compliance
  4. Organizational/Administrative
  5. Cultural/Professional Norms
  6. Resource/Staffing

Completed Fishbone (Summary):

CATEGORY 1: Clinical Process

  • Main: Documentation required during visit (not pre/post) itemize
  • Sub: Billing codes require specific data captured in real-time
  • Sub: Quality metrics tied to in-visit documentation completion itemize
  • Main: Complex visit workflows (multiple chronic conditions per patient) itemize
  • Sub: Average 3.2 chronic conditions per visit (observed)
  • Sub: Each condition requires separate documentation itemize
  • Main: No standardized documentation templates by visit type
  • Evidence: Observation—all 12 observed visits required real-time documentation; Interview P05 (physician): "Can't document after—will forget details for billing"
  • Addressability: A (addressable)—workflow redesign, pre-visit prep, templates

CATEGORY 2: Technology/EMR

  • Main: EMR requires excessive click burden (15-20 clicks for common order) itemize
  • Sub: No smart defaults—must specify every parameter manually
  • Sub: Frequent medication not in favorites (search required) itemize
  • Main: Interface not optimized for in-visit use (small text, dense screens)
  • Main: No voice input capability (must type)
  • Main: Alerts/interruptions during documentation (pop-ups break flow)
  • Evidence: Observed—P08 took 14 clicks to order routine lab; Interview P11: "Alerts constantly interrupt, have to dismiss 10 pop-ups per visit"
  • Addressability: A (addressable)—EMR optimization, voice input integration, alert reduction

CATEGORY 3: Regulatory/Compliance

  • Main: Meaningful Use requirements mandate specific data fields
  • Main: MIPS quality reporting requires discrete data elements (not narrative)
  • Main: Billing/coding rules require contemporaneous documentation
  • Evidence: Interview P02 (compliance): "If not documented in structured fields during visit, doesn't count for MIPS"; Policy documents confirm contemporaneous requirement
  • Addressability: C (constraint)—regulatory requirements outside clinic control, must accept

CATEGORY 4: Organizational/Administrative

  • Main: Visit time not adjusted for documentation burden (same 15-min slots as pre-EMR)
  • Main: No scribe support (budget constraints)
  • Main: Productivity metrics incentivize volume over visit quality
  • Main: No protected time for post-visit documentation ("pajama time" at home)
  • Evidence: Interview P06: "Still scheduled 15 min visits even though now spend 7 min on computer"; P14: "Do 2 hours documentation at home after dinner—pajama time"
  • Addressability: Mixed—D (deferred) for scribe/scheduling (clinic administration decision), A (addressable) for documentation efficiency tools that reduce time burden

CATEGORY 5: Cultural/Professional Norms

  • Main: Physicians expected to personally document (delegation not acceptable)
  • Main: Narrative clinical reasoning valued over structured data (professional identity)
  • Main: Resistance to standardized templates ("cookbook medicine" stigma)
  • Evidence: Interview P09: "Real doctors write thoughtful notes, not check boxes"; P12: "Templates feel like I'm not thinking"
  • Addressability: C (constraint)—professional culture outside product scope, but can design to respect narrative while capturing structured data

CATEGORY 6: Resource/Staffing

  • Main: Physician shortage → high patient load → time pressure
  • Main: No dedicated IT support for EMR optimization
  • Main: Insufficient training on EMR efficiency features (physicians self-taught)
  • Evidence: Observation—physicians unaware of shortcuts, favorites, templates available in EMR; Interview P07: "Never trained on efficient use, just trial and error"
  • Addressability: A (addressable)—training program, efficiency coaching, A (partial)—IT support advocacy

Cross-Category Patterns Identified:

  • Interactive effect: Regulatory requirements (Category 3) + EMR poor design (Category 2) + time pressure (Category 6) compound—each alone manageable, combination creates burden
  • Root cause of root causes: Visit time not adjusted (Category 4) amplifies all other categories—if had adequate time, click burden and complexity more tolerable

Prioritization (Addressable Causes):

Top 3 addressable causes (dot voting result):

  1. EMR click burden and alert overload (Category 2)—18 votes
  2. No standardized templates by visit type (Category 1)—15 votes
  3. Insufficient EMR efficiency training (Category 6)—12 votes

Boundary Definition (Feeding Step 4):

In-scope (addressable):

  • EMR interface optimization (reduce clicks, smart defaults, alert reduction)
  • Visit-type-specific documentation templates
  • Voice input integration (if technically feasible)
  • Physician efficiency training program

Out-of-scope (constraints accepted):

  • Regulatory requirements (Meaningful Use, MIPS, billing rules)—design must comply
  • Professional culture (narrative preference)—design must accommodate, not fight
  • Physician shortage (systemic healthcare issue)—solution must work within time constraints

Deferred (future or parallel initiatives):

  • Scribe program (administrative decision, separate budget)
  • Visit scheduling redesign (clinic operations, separate workstream)
  • Post-visit documentation time allocation (HR/compensation policy)

Solution Implication:

Ishikawa revealed that problem is multi-causal—not solvable with single intervention. Solution must address:

  1. Technology optimization (reduce click burden)
  2. Workflow redesign (templates, pre-visit prep)
  3. Training (efficiency techniques)

Attempting only (1) without (2), (3) would yield <30% improvement. Comprehensive approach addressing top causes yields 60-70% documentation time reduction (validated in pilot).

If team had used Five Whys instead: Would likely have followed single path (e.g., "Why burden? → EMR poor design → Why poor? → Vendor prioritizes billing over UX → Why vendor? → Procurement decisions"). Would miss that problem is systemic, not just EMR.

Ishikawa's multi-category structure revealed true complexity and guided comprehensive solution.

Journey Mapping method:journey-mapping

Used in:A1.2 Step 2(for temporal/process needs)

[Complete 2,500-word content with interview/workshop approaches...]

Need Assessment Rubrics method:need-assessment

Used in:A1.2 Step 5(Assessment)

[Complete 2,000-word content with 5-dimension framework...]

Need-Focused Interviewing method:interviewing

Used in:

  • A1.2 Step 2.1 (Need discovery)
  • A1.3 Step 2 (Root cause exploration)
  • A1.4 Step 3 (Multi-dimensional understanding)
  • A6 Step 3 (Concept validation)

Core Method: Qualitative inquiry across discovery and validation

Purpose: Conduct structured conversational inquiry to deeply understand user experiences, needs, motivations, and reactions—adapted to specific investigation purpose (need discovery vs. root cause vs. concept validation).

Processing Workflow Overview

A1.2 data processing follows a systematic workflow from raw data to synthesized insights:

figure[H] tikzpicture[ node distance=1.3cm, process/.style= rectangle, rounded corners, minimum width=3.5cm, minimum height=0.9cm, text centered, draw=alphacolor, fill=alphacolor!15, font= , arrow/.style= thick, ->, >=stealth, color=alphacolor ]

(step1) [process] 1. Transcript & Note Preparation; (step2) [process, below=of step1] 2. Initial Coding & Insight Extraction; (step3) [process, below=of step2] 3. Affinity Diagramming (Thematic Clustering); (step4) [process, below=of step3] 4. User Segmentation (2x2 Matrix); (step5) [process, below=of step4] 5. Synthesis Artifacts Creation; (step6) [process, below=of step5] 6. Quality Review & Validation;

[arrow] (step1) – (step2); [arrow] (step2) – (step3); [arrow] (step3) – (step4); [arrow] (step4) – (step5); [arrow] (step5) – (step6);

at (5.5, -0.5) [font=, text width=3cm] Clean, organize, anonymize; at (5.5, -2) [font=, text width=3cm] Extract atomic insights from data; at (5.5, -3.5) [font=, text width=3cm] Group insights into themes; at (5.5, -5) [font=, text width=3cm] Create user typology; at (5.5, -6.5) [font=, text width=3cm] Journey maps, personas, empathy maps; at (5.5, -8) [font=, text width=3cm] Verify rigor & completeness;

tikzpicture Data Processing Workflow - 6 Steps from Raw Data to Synthesis fig:data-processing-workflow-appendix figure

Typical timeline: 10-17 days (2-3 weeks) for comprehensive synthesis after data collection completes.

Team requirement: All researchers who conducted interviews should participate in synthesis—direct exposure to users is essential for accurate interpretation.

Critical distinction maintained throughout this workflow:

  • Clustering/Affinity Diagramming (Step 3) = Data synthesis method organizing interview/observation insights into thematic groups
  • 2x2 Matrix Segmentation (Step 4) = User typology creation method segmenting users by psychographic dimensions

These are different methods serving different purposes—clustering synthesizes what was said, segmentation categorizes who said it and how they differ.

When to Use:

  • Users can consciously articulate experiences (vs. latent/tacit needs requiring observation)
  • Need deep exploration of motivations, emotions, decision-making, impacts
  • Want to understand “why” behind behaviors (complement to observation's “what”)
  • Flexible method applicable across discovery, definition, and validation stages

Core Interview Structure (60-75 minutes):

[Detailed protocol - 5,000 words covering all aspects]

  • Opening and rapport building (5 min)
  • Context setting (10 min)
  • Core inquiry (35-40 min) [Varies by activity—see adaptations below]
  • Impact and reflection (10 min)
  • Closing and referrals (5 min)
  • H. Beyer & K. Holtzblatt (1997). Contextual Design: Defining Customer-Centered Systems. Morgan Kaufmann.
Coming soon

Share how you use Affinity Diagramming

This is where practitioners will be able to share field notes, variations, and additional templates for this method — what worked, what to watch for, and adaptations for different contexts.

Until the community space opens, we welcome contributions by email and will fold the best into the method page.