◀ Stage Select World 4 · Stage 4-01

Cohort Analysis

Stop averaging away the truth

A cohort is a group of users who started together. Tracking each group separately as it ages shows whether the product is actually getting better, something a blended, all-users average can hide completely. This guide covers how to build a cohort view, read it, and avoid the traps that make it lie.

Ready?

Step-by-step lessons

See What the Average Hides

Four short lessons: what a cohort is and why averages mislead, how to read a cohort table, the cohort types and what each answers, and the traps that quietly break the analysis.

1

What a Cohort Is, and Why Averages Lie

A cohort is a group of users who share a starting event inside the same time window, tracked together as the group ages. Sign-up week is the usual starting event. The point of grouping this way is simple: everyone in a cohort has been a user for the same length of time, so you can compare like with like.

A blended, all-users number can't do that. It mixes people on day 3 with people on day 900, so it moves whenever the mix moves, even if nothing about the product changed.

Everyday example, a gym in January

A gym's average member attendance jumps every January. Not because members got fitter, but because a flood of new joiners temporarily changes who "the average member" is. By March the average falls again, and still nothing about the gym has changed. Track the January joiners as their own group and you learn something real: how many of them are still turning up in week 6.

Blended number

"Retention is 41%"

41% of whom, measured how long after what? Moves with acquisition volume, not product quality.

Cohort number

"March signups, 41% active at week 4"

A specific group, a specific age, a stated definition of active. Comparable to April's week 4.

Quick check

A company's blended retention number has risen for three straight months, but acquisition has slowed sharply over the same period. What's the most likely explanation?

2

Reading a Cohort Table

A cohort table is a triangle. Each row is one cohort, each column is an age (week 1, week 2, week 3…), each cell is the share of that cohort still active at that age. It's a triangle because the newest cohort hasn't lived long enough to fill the later columns yet.

Read across a row

One cohort's decay curve. Steep early, then hopefully flattening.

Read down a column

Different cohorts at the same age. This is the honest read on whether the product is improving.

Read the diagonal

One calendar date hitting every cohort at once, an outage, a holiday, a bad release, a tracking break.

Check the size column

A 60% cell on 20 users means nothing. Percentages need a denominator worth trusting.

The shape that matters most

Every retention curve falls at first. The question is where it stops falling. A curve that flattens at 30% says a real group found lasting value, and that 30% is the foundation the business compounds on. A curve that keeps sliding toward zero is a leaky bucket: growth is being rented from the acquisition budget, and it stops the moment spending does.

You're not looking for a high number. You're looking for a curve that goes flat above zero.

Quick check

Every cohort in a table dips sharply in the same calendar week, regardless of how old each cohort was at the time. How should this be read?

3

Cohort Types, and What Each One Answers

"Cohort" isn't one thing. How you group people decides which question the table can answer.

By start date

Acquisition cohort

Grouped by signup week or month. Answers: is the product getting better for each new intake?

By action

Behavioral cohort

Grouped by an action taken in a window, "connected a calendar in week 1". Answers: which early behavior predicts sticking around?

By source

Channel or campaign cohort

Grouped by where users came from. Answers: which acquisition source brings users who stay?

By value

Revenue cohort

Tracks money, not logins. Net revenue retention can exceed 100% when expansion beats churn.

Two definitions to pin down before anyone reads a number. Active: opened the app, or completed a core action? The second is harder and far more honest. Retained at week N: active in week N specifically (bracket retention), or active at any point up to week N (unbounded)? Unbounded always looks better and is rarely the number you want.

Behavioral cohorts show correlation

"Users who connect a calendar retain 3× better" is a great hypothesis and a terrible conclusion. Those users may simply have been more committed to begin with. Forcing the action on everyone can add friction and move nothing. Behavioral cohorts nominate an activation event; an experiment is what promotes it to a cause.

Quick check

A B2B company reports 92% logo retention at month 12, but net revenue retention of 118% for the same cohort. How is that possible?

4

Traps, and Turning the Table into a Decision

Cohort tables fail quietly. Five traps account for most bad readings.

1

Immature cohorts

The newest rows have partly-elapsed periods. Those trailing cells are censored, not bad. Grey them out.

2

Tiny cohorts

Forty-user weekly cohorts swing 20 points on a handful of people. Widen the grain or pool weeks.

3

Changed definitions

Redefining "active" mid-table breaks every column comparison unless the raw events allow a clean backfill.

4

Mixed channels

A flat aggregate can be one channel collapsing while another improves. Segment before concluding.

5

Survivorship bias

Studying only the users who stayed hides everyone who did the same thing and left anyway.

Then act. Attack the earliest steep part of the curve first: users lost in week 1 are the cheapest to save and the largest group. Compare your best and worst cohorts and ask what differed, the channel, the onboarding version, the season. And when a column moves after a release, you have a lead worth testing properly.

Quick check

Every cohort's retention decays toward zero by month 9, yet the company's total active users keep climbing. What does this combination mean?

Review the concepts

Cohort Analysis Flashcards

6 cards covering the essentials. Click a card to flip it.

Click to flip

Tip: say the answer out loud before flipping.

Explanation

In practice

1 / 6
Did you recall it?
Apply what you learned

Practice Scenarios

15 situations that test whether you can define a cohort, read a retention table, and avoid the traps that make cohort data lie.

Scenario 1easy

A PM reports "our retention is 41%" with no other qualifiers attached to the number.

What's missing from that statement?

Scenario 2easy

A team's blended, all-users retention number has been climbing for three months, and leadership is celebrating.

Why might cohort analysis tell a different story?

Scenario 3easy

A team groups all users who signed up in March 2026 and tracks that group's behavior over the following six months.

What kind of cohort is this?

Scenario 4easy

A cohort retention curve drops steeply for the first two weeks, then levels off at around 30% and stays flat for months.

What does the flattening suggest?

Scenario 5easy

Someone builds a cohort table and compares the newest cohort's month-1 retention against a year-old cohort's month-12 retention.

What's wrong with this comparison?

Scenario 6medium

Reading down the month-1 column of a cohort table, retention goes 22%, 24%, 27%, 31% across four consecutive monthly cohorts.

What does reading down a column tell you?

Scenario 7medium

In a cohort table, every cohort shows an unusual dip in the same calendar week, regardless of the cohort's age.

How should the team read this pattern?

Scenario 8medium

A subscription business reports 92% logo retention at month 12, but its month-12 net revenue retention is 118%.

How can revenue retention exceed 100%?

Scenario 9medium

A PM defines a behavioral cohort of "users who connected a calendar in week 1" and finds they retain at triple the rate of everyone else.

What should the PM conclude?

Scenario 10medium

A monthly cohort table shows the two most recent cohorts performing far worse than all previous ones at month 3.

What should the team check first?

Scenario 11hard

A weekly cohort table is sliced by acquisition channel. Overall retention is flat, but paid-social cohorts are collapsing while organic cohorts are improving sharply.

What's the real story?

Scenario 12hard

Mid-year, a team changes the definition of an "active user" from opening the app to completing a core action, and backfills the new definition across the whole cohort table.

What's the risk to the analysis?

Scenario 13hard

A team runs cohort analysis on weekly cohorts of about 40 users each, and the week-4 retention numbers swing between 18% and 47% with no clear trend.

What's the most likely explanation?

Scenario 14hard

A PM studies the users still active at month 12 to find what drove their retention, and finds they all used the advanced reporting feature heavily.

What bias threatens this conclusion?

Scenario 15hard

A cohort curve shows retention decaying steadily toward zero by month 9 for every cohort, but the company's total active users keep growing.

What does this combination indicate?

Lock it in

Guess the Term

Read the clues and name the concept. The fewer clues you need, the more points you score.

Round 1 Score 0
Keep it handy

Cohort Analysis Quick Reference

The whole topic on one screen.

Reading the Table

Row

One cohort's decay as it ages.

Column

Cohorts compared at the same age.

Diagonal

One calendar date hitting everyone.

Cohort Types

Acquisition

Grouped by signup date.

Behavioral

Grouped by an early action.

Channel

Grouped by acquisition source.

Revenue

NRR can exceed 100% via expansion.

Curve Shapes

Flattens above zero

A stable core found lasting value.

Decays to zero

Leaky bucket, growth is rented.

Column rising

Newer cohorts doing better at the same age.

Traps

Immature cohorts

Partial periods look like collapse.

Small samples

Tiny cohorts swing on a few users.

Changed definition

Breaks comparability across rows.

Survivorship bias

The churned users are invisible.

Notification