Team health checks: how to measure how your team is really doing
What a team health check is, when to run one, which questions to ask by dimension, and how to turn the scores into change your team can see. With an honest read on the Spotify and Atlassian models.
A team health check is a short, structured self-assessment in which everyone on a team anonymously rates the same set of dimensions, such as delivery, ownership, and psychological safety, on a simple scale, then discusses the spread together. Run on a regular cadence, it turns “how are we doing?” from a guess into a trend you can act on.
That is the whole idea. The rest of this guide is the practice: how a health check differs from a retrospective and an engagement survey, when to run one, how to choose dimensions (including an honest look at the Spotify and Atlassian models), a question bank organized by dimension, how to facilitate the session so people answer truthfully, and what to do with the results once you have more than one of them.
What a team health check is (and what it isn’t)
Two other rituals get confused with health checks, and the confusion matters because each one answers a different question.
A health check is not a retrospective. A retrospective examines a specific slice of work: what happened last sprint, what to change next sprint. Its raw material is recent events, and its output is actions. A health check measures the team itself against a stable set of dimensions, and its output is scores you can compare quarter to quarter. A retro tells you what happened; a health check tells you what it is like to work here, and whether that is getting better or worse. The two pair naturally — retro every sprint, health check every quarter, often run inside a retro slot.
A health check is not an engagement survey. Engagement surveys are owned by HR, run across the whole organization once or twice a year, and ask about the individual’s experience of the company. The results travel up. A health check is owned by the team, runs often, asks about the team’s way of working, and the results stay in the room where they can be acted on. Both are legitimate instruments. The trouble starts when an annual survey is treated as a substitute for a team knowing its own state, or when a health check gets repurposed as a management scoreboard (more on that below).
| Retrospective | Team health check | Engagement survey | |
|---|---|---|---|
| Looks at | The last sprint or period of work | The team’s way of working | The individual’s experience of the org |
| Owned by | The team | The team | HR / leadership |
| Cadence | Every sprint | Quarterly (typically) | Annually or twice a year |
| Output | Actions | Scores, trends, and a conversation | Org-level metrics |
| Results go | Into the team’s next sprint | Back to the team | Up the reporting line |
When to run one, and how often
Quarterly is the right default. It is frequent enough to catch drift while it is still cheap to fix, and infrequent enough that the scores have a chance to move between checks. Running the same check every sprint mostly measures noise; running it annually measures a team that no longer exists.
Beyond the regular beat, run an extra check whenever the team changes shape: after a reorganization, when a new lead arrives, when several people join or leave at once, or when a team is being formed from scratch. Those are exactly the moments when everyone’s private read on the team diverges most, and a health check makes the divergence visible early.
Two variations on cadence are worth knowing. Teams in active turnaround sometimes move to a monthly pulse on a reduced question set, then drop back to quarterly once the trend stabilizes. And distributed teams often run the rating step asynchronously over a couple of days, then meet only for the discussion — the scores wait patiently; the conversation is the part that needs everyone present.
Choosing your dimensions
A dimension is one aspect of team health you will rate every time: “delivery”, “psychological safety”, “workload”. Choosing them is the most consequential design decision you will make, because the dimensions define what the team considers worth measuring — and because you need to keep them stable across quarters or the trend line means nothing.
Two published models dominate this space, and both are worth learning from honestly.
The Spotify squad health check model
Published by Henrik Kniberg and Kristian Lindwall in 2014, the Spotify squad health check model has squads rate around eleven dimensions (delivering value, easy to release, health of codebase, teamwork, mission, and so on) as green, yellow, or red in a facilitated workshop, producing a colored grid across squads.
What it is genuinely good at: it made team health visible at organizational scale, and its traffic-light simplicity gets a real conversation going fast. It reframed measurement as self-assessment for the team’s benefit rather than reporting for management, and a decade later that framing still holds up.
Its honest limitations: the dimensions were designed for Spotify’s engineering context, and several (health of codebase, easy to release) map poorly onto non-engineering teams. Spotify’s own authors cautioned against copying their practices wholesale. And a three-color scale is coarse — good for a workshop conversation, weak for detecting a slow six-month slide from “strong green” to “barely green”.
Atlassian’s Team Health Monitor
Atlassian’s Health Monitor takes a different angle: eight attributes of a healthy team — among them a balanced team, shared understanding, and suitable ways of working — self-scored as a group in a facilitated workshop. (Atlassian has since consolidated its older project-, service-, and leadership-team versions into one shared attribute set.) Its strength is structural clarity: attributes like these surface role and alignment problems that mood-centric checks miss entirely, and it works well outside software.
Its limitations are the mirror image: it is a workshop instrument, scored openly in discussion rather than anonymously, so it is only as honest as the room feels safe. And as a point-in-time exercise it gives you a snapshot, not a trend, unless you impose your own cadence and record-keeping around it.
Our recommendation
Take the framing from both models and fix the two weaknesses: rate anonymously, and rate on a scale you can trend. Choose six to ten dimensions, mix delivery-facing ones (process, roles, value) with human ones (safety, workload, collaboration), adapt the wording to your context — and then stop editing them. A dimension set that changes every quarter is a new check each time, and you lose the only thing that compounds: the trend. If you would rather not design from scratch, start from a ready-made team health check template and trim it.
Those two fixes — anonymous rating and a scale you can trend — are also the right lens for choosing a tool to run this in, because most options get one or the other wrong. We compare the honest choices, from dedicated tools to the Spotify framework to HR-survey platforms, in our roundup of the best team health check tools.
Team health check questions, by dimension
Ask people to rate statements, not to answer open questions. A statement like “I can raise a problem without worrying how it lands” produces a score you can compare across quarters and a concrete conversation starter in one line. Rate each on a five-point agree–disagree scale.
Direction and value
- We know what we are trying to achieve and why it matters.
- Our priorities are clear enough that we can say no to work that doesn’t fit.
- I can connect what we shipped recently to value for users or the business.
Delivery and process
- Our way of working helps us more than it slows us down.
- Releasing or handing over work is routine, not an event.
- Work rarely sits stuck waiting on another team or an approval.
Roles and ownership
- I know what is mine to decide and what needs the group.
- Important decisions have a clear owner.
- When something falls between roles, we notice and assign it quickly.
Psychological safety and trust
- I can raise a problem without worrying how it lands.
- Mistakes here lead to fixes, not blame.
- Disagreement on this team is treated as useful information.
Collaboration and communication
- I know enough about what others are working on to spot overlaps and gaps.
- When someone asks for help, help arrives.
- Our meetings earn the time they take.
Learning and improvement
- We change how we work based on what we learn.
- The actions we agree in retrospectives actually get done.
- Time for learning survives our busy periods.
Workload and sustainability
- The current pace is one we could keep up for a year.
- I can finish the important things without routinely working evenings.
- Interrupt work and on-call load are shared fairly.
Support and resources
- We have the tools and access we need to do the job well.
- When we escalate a blocker, someone above us acts on it.
- The team’s case is heard when priorities are set above us.
Trim rather than add. Twenty-four statements is the ceiling, not the target — a check people can finish thoughtfully in five minutes beats a thorough one they start skimming.
Psychological safety deserves its own reading
One dimension consistently predicts the usefulness of everything else: psychological safety, the shared belief that this team is safe for interpersonal risk-taking. Amy Edmondson’s research established the concept, and Google’s Project Aristotle later found it the single strongest differentiator among its own effective teams. It also has a compounding property inside a health check: a team that scores low on safety is probably inflating its other scores too, because the check itself is an act of speaking up. Our full guide to psychological safety at work covers how to build and measure it as a standalone practice.
A useful lens for interpreting a safety score is Timothy R. Clark’s 4 stages of psychological safety: inclusion safety (I belong here), learner safety (I can ask questions and make mistakes), contributor safety (I am trusted to do real work), and challenger safety (I can question how things are done). The stages are a ladder, and most teams sit further up it than they think — plenty of teams where everyone feels included would still go quiet if someone questioned the roadmap. If your safety scores are middling, the follow-up question is which stage is missing, and Clark’s labels give the team language for that conversation.
Be honest about what a health check can and cannot do here. Three rated statements will tell you whether safety needs attention; they will not diagnose it. If the scores say there is a problem, that finding deserves its own session — a dedicated psychological safety check goes deeper than a general health check should, and our team dynamics guide covers how trust and safety are actually built.
How to run the session
The mechanics matter more than they look. The same questions, run badly, produce polite fiction.
- Make ratings anonymous. This is the non-negotiable one. The moment people sign their names to a low score, the low scores dry up, and you are measuring what people are willing to say rather than what they think. Anonymity is also why a tool beats fists-of-five in a meeting: nobody can watch the room before choosing their number.
- Rate first, discuss second. Everyone scores every statement silently before any discussion. If the conversation happens first, the ratings anchor on whoever spoke, and usually on whoever is most senior.
- Discuss the spread, not the average. A dimension where everyone scored 3 and a dimension where half the team scored 5 and half scored 1 have the same average and mean completely different things. The disagreement is the interesting data: it means two groups of people are having different experiences of the same team.
- Timebox to under an hour. Five minutes of rating, then the rest on the two or three dimensions most worth talking about. You are not obliged to discuss every dimension every time; that is what the trend line is for.
- Leave with one or two owned actions. Pick the dimension you most want to move, agree one change, and give it a named owner and a date. Our Follow-Through Index data is blunt on this point: actions with an owner and a due date get completed about 90% of the time, and actions with neither mostly don’t.
What to do with the results over time
A single health check is a conversation. The value compounds when you keep the questions stable and run it on a rhythm, because then you have a trend — the same team radar of dimensions, read across quarters instead of once. A trend answers questions a snapshot cannot. Is the workload problem seasonal or structural? Has safety recovered since the reorganization?
Three habits make the trend worth having:
- Open each check by reviewing the last one. What did we say we would change, did we change it, and did the dimension move? A health check whose actions are never revisited teaches the team that scoring honestly changes nothing, and the scores will quietly converge on a safe, meaningless 4.
- Move one dimension at a time. A team that tries to improve six dimensions at once improves none of them. Pick the one that matters most this quarter and put real effort behind it; let the others ride.
- Use cross-team views to direct support, not to rank teams. Seeing health across many teams is genuinely useful for a coach or a leadership group deciding where help is needed. The moment scores become a league table, teams start managing the number instead of the work, and the instrument is dead. Share trends, not comparisons, and never tie scores to performance reviews.
Maturity models: the zoomed-out cousin
Health checks have a sibling worth knowing about. A health check asks “how does it feel to work here right now?” — subjective by design, and answered by the whole team every quarter. A maturity model asks a different question: “how developed are our practices against a defined scale?” It assesses concrete practices (how you plan, test, release, measure) against staged levels, and it is typically revisited every six to twelve months rather than quarterly.
The two complement each other. The health check is the thermometer; the maturity model is the map. A team can be happy and immature (great atmosphere, chaotic releases) or mature and miserable (disciplined practices, unsustainable pace), and you only see the full picture with both readings. If your health-check trend keeps flagging the same delivery dimensions, that is often the cue to zoom out: our agile maturity assessment walks a team through exactly that practice-level review, and the accompanying guide to measuring and improving agile maturity covers how to act on what it finds.
Templates to start from
You can design a health check from a blank page, but you rarely need to. These ready-to-run templates cover the common starting points, and every one of them runs with anonymous rating and trend tracking built in:
- Team health check — the balanced general default across delivery, collaboration, and wellbeing.
- Squad health check — the Spotify model, ready to run.
- Psychological safety check — a deeper standalone read on safety and trust.
- Team effectiveness health check — focused on how well the team converts effort into outcomes.
- Team dysfunction radar — based on Lencioni’s five dysfunctions, for teams that suspect a deeper pattern.
Or browse the full health check template library for role-specific and situation-specific variants.
Ready to take the guesswork out of “how are we doing?” TeamRetro’s health checks run the whole loop described in this guide — anonymous ratings, dimension-by-dimension discussion, trends across quarters, and actions tracked through to done. Start a free trial and run your first check this week.
Keep reading
- The Follow-Through Index — our first-party data on whether teams do what they decide.
- Team dynamics: how great teams actually work — the forces your health check is measuring.
- Running effective retrospectives — the sprint-level companion ritual.
Frequently asked questions
What is a team health check?
A team health check is a short, structured self-assessment in which everyone on a team anonymously rates the same set of dimensions, such as delivery, ownership, and psychological safety, on a simple scale, then discusses the spread together. Run on a regular cadence, it turns “how are we doing?” from a guess into a trend you can act on.
What is the difference between a team health check and a retrospective?
A retrospective examines a specific slice of work, usually the last sprint, and produces actions. A health check measures the team itself against a stable set of dimensions and produces scores you can trend over time. They pair well: retro every sprint, health check every quarter, often run inside a retro slot.
How often should you run a team health check?
Quarterly is the right default for most teams: frequent enough to catch drift, infrequent enough that scores have a chance to move. Add an extra check when the team changes shape, such as after a reorganization, a new lead, or several new joiners. Monthly pulse checks work for teams in active turnaround; weekly is survey fatigue.
What questions should a team health check ask?
Ask people to rate statements, not answer open questions, and organize them by dimension: direction and value, delivery, roles and ownership, psychological safety, collaboration, learning and improvement, workload, and support. Statements like “I can raise a problem without worrying how it lands” or “the current pace is one we could keep for a year” give you a comparable score and a conversation starter in one line.
What are the 4 stages of psychological safety?
Timothy R. Clark’s model describes four stages a person moves through on a team: inclusion safety (I belong here), learner safety (I can ask questions and make mistakes), contributor safety (I can do real work and be trusted with it), and challenger safety (I can question how things are done). It is a useful ladder for reading a team’s health-check scores: a team can feel included yet still be unsafe to challenge.