# Hypothesis Retrospective (QA Focused)

Cross-functional teams surface quality bottlenecks and convert them into structured, measurable experiment statements for their next sprint.

**Time:** 55 minutes. UWT suggested 55-minute duration. The source mentions a 10-minute brainstorming round followed by grouping, voting, and hypothesis drafting.
**People:** 4 to 24. Solo silent writing, followed by small cross-functional teams of 3 to 5 people (up to 5 groups total), ending in whole-group voting.

## Choose this exercise

**Purpose:** Convert recurring product defects and testing bottlenecks into concrete, measurable improvement experiments.
**Use it when:** Use at the end of a sprint or release when quality issues slow down delivery or cause friction between developers, testers, and product managers.
**Skip it when:** Skip during an active production incident or when the team lacks any baseline delivery data to track against.

## Materials and tools

- Sticky notes (three pads per small group)
- Fine-tip markers (one per participant)
- Dot stickers (6 per participant)
- Printed copies of the Hypothesis Construction Template (one per small group)
- Whiteboard or wall space with 4 labeled hypothesis template columns: 'We hypothesize by', 'We will', 'Which will', 'As measured by'
- Shared digital whiteboard platform and digital dot tokens (for remote or hybrid delivery)

## Before participants arrive

1. Draw the four-part hypothesis canvas on a central board: 1. We hypothesize by [implementing this], 2. We will [improve on this problem], 3. Which will [create this benefit], 4. As measured by [this metric].
2. Prepare a dedicated staging area to the side for raw defect and bottleneck notes.

## Say this to open

Welcome. Quality is a whole-team discipline, but testers often spot systemic defects that development misses. Today we will surface our biggest quality bottlenecks and shape them into measurable experiments. First, everyone will silently write down the quality issues they saw during this cycle.

## How to run it

### 1. Brainstorm quality bottlenecks (10 minutes)

Open with the opening script. Participants write one quality issue or defect pattern per sticky note in silence. Examples include late testing handoffs, missing test data, or ambiguous acceptance criteria. Participants post their notes in the staging area as they finish.

### 2. Cluster and select top issues (15 minutes)

Participants step up to the board and silently group related notes into theme clusters for 6 minutes. The facilitator resolves overlaps and writes a clear header above each cluster for 4 minutes. Give each participant 3 dot stickers. Participants silently place dots on the clusters causing the greatest quality drag. The top clusters advance, matching the number of small teams to be formed (2 to 5 clusters).

### 3. Draft testable hypotheses (15 minutes)

Form cross-functional teams of 3 to 5 people. Ensure developers and testers are distributed evenly. Assign one top issue cluster to each team. Teams complete the Hypothesis Construction Template using its four prompts. Example: We hypothesize by requiring smoke tests on pull requests / We will reduce broken main builds / Which will give QA stable test environments / As measured by zero build-break alerts next sprint.

### 4. Vote and assign commitment (15 minutes)

Each team posts their hypothesis sheet and presents it in 60 seconds. Participants place 3 dot stickers on the experiment they believe offers the highest impact. Tally the votes to identify the top hypothesis. Facilitate a brief debrief with up to 2 questions and 3 participant responses. Confirm one owner and log the experiment into the sprint backlog.


## Debrief questions

- Which metric in our chosen hypothesis will be simplest to verify during daily work?
- What specific support does engineering or product need to run this experiment cleanly?

**Finished output:** One prioritized hypothesis statement with an assigned owner, clear metric, and backlog commitment for the next sprint.

## Remote

Set up a digital whiteboard with a staging area and breakout frames containing the four-part template. Run steps 1 and 2 in the main room. Send small groups of 3 to 5 to breakout rooms for step 3. Reconvene in the main room and use digital dot tokens to vote on hypotheses in step 4.

## Hybrid

All participants use a single central digital whiteboard for posting, clustering, and voting. In-person participants work at physical tables or in pods using a room microphone and webcam. If there are 3 or more remote participants, they form their own remote breakout pod; otherwise, individual remote members join in-person table pods with an assigned laptop. Clusters are assigned on the shared digital board so all pods see their task simultaneously.

## Access and participation choices

- Participants may write notes silently or dictate them to a peer or the facilitator, who writes and posts the sticky note.
- Participants can submit issue cards anonymously before the session.
- Any participant may pass on vocal presentation by designating a peer to read the group's hypothesis.

## If the session gets stuck

**Teams propose broad habits like 'communicate earlier' rather than testable practices.**

Ask: 'What specific checklist, meeting, or automated step can we add or remove next week?'

**Groups struggle to identify an objective metric.**

Guide them toward countable events, such as 'fewer than 2 escaped bugs' or 'zero blocked staging deployments'.


## Terms used here

- QA: Quality Assurance, the discipline and team role focused on identifying defects and testing software behavior.
- Dot-voting: A quick prioritization method where participants place dot stickers or digital tokens on options.

## Participant material: Hypothesis Construction Template

Problem Cluster:
What quality defect, bottleneck, or friction point are we tackling?

Hypothesis Frame:
1. We hypothesize by: [What specific change or practice will we adopt?]
2. We will: [What specific issue will this alter or improve?]
3. Which will: [What direct positive outcome will the team see?]
4. As measured by: [What observable count, log, or signal will prove it worked?]

Commitment:
Review date: [End of next cycle]
Assigned owner: [Name]

## Sources and adaptation

Observed on EasyRetro as a retrospective template attributed to Quality Logic.

Retained the four-part hypothesis template ('We hypothesize by', 'We will', 'Which will', 'As measured by') and the core sequence of individual issue generation, clustering, voting, and small group hypothesis drafting from EasyRetro and Quality Logic. UWT structured the agenda into a realistic 55-minute format, balanced group math to accommodate 3 to 5 people per team across any roster from 4 to 24, specified silent clustering with facilitator support, and added detailed hybrid, access, and facilitation steps.

[EasyRetro](https://easyretro.io/templates/hypothesis-retrospective-qa-focused/)

Current run sheet: https://unitedwetransform.com/exercises/hypothesis-retrospective-qa-focused/
