Cohort Retention Analysis: How to Build It and What It Reveals
Blended churn averages good cohorts with bad ones and hides the trend that matters. How to build a cohort retention triangle, where the curve flattens, and why revenue cohorts can exceed 100%.
Key takeaways
- A single churn number mixes every cohort together and tells you nothing about whether retention is improving or deteriorating.
- Cohort retention groups customers by the period they first paid, then tracks what that same group spends over time, which is the only way to see direction.
- The month the retention curve flattens is your genuinely retained base, and it is the only defensible input to a lifetime value calculation.
- Revenue cohorts can exceed 100% because existing customers expand. That is the same effect that lets net revenue retention exceed 100%, and it is a good sign.
- Rapid acquisition can hold total revenue flat while the existing base erodes underneath. Only a cohort view separates the two.
Cohort retention analysis groups customers by the period they first paid you, then tracks what that same group spends in every subsequent period. It exists because a single blended churn number averages good cohorts with bad ones and tells you nothing about direction, and direction is the only thing that lets you know whether what you changed last year worked.
What a blended churn rate hides
Suppose your annual gross churn is 12%, and it was 12% last year. That looks like stability. Underneath it, either of these could be true:
- Every cohort churns at a steady 12%, and nothing has changed.
- Cohorts from two years ago now churn at 20% while cohorts acquired this year churn at 6%: the product got dramatically better, and the blended number is being dragged by a legacy base that is shrinking anyway.
These are opposite situations with identical headline metrics. The second is a company whose fundamentals improved; the first is a company standing still. Only a cohort view distinguishes them.
Building the triangle
The data requirement is modest: one row per customer per period, with a revenue amount. Three columns.
- Assign cohorts. Each customer's cohort is the first period in which they appear.
- Calculate age. For every row, the number of periods elapsed since that customer's cohort period. Month 0 is their first period.
- Aggregate. Sum revenue by cohort and age. This gives you the triangle: cohorts down the side, age across the top.
- Convert to retention. Divide each cell by that cohort's month-0 value.
The shape is triangular because newer cohorts have not existed long enough to fill the later columns. That is expected, and it is also the main way people misread the chart: comparing a column across cohorts is valid, comparing cohorts at different ages is not.
In a spreadsheet this means a pivot table, a triangular layout, and array formulas that break whenever a customer skips a period. It is genuinely fiddly, which is why most teams build it once, discover an error, and go back to reporting blended churn.
Reading the curve
Where it flattens is your real retained base
Most retention curves drop steeply for the first few periods and then flatten. The customers who were a poor fit leave early; the ones who remain tend to stay.
The month the curve flattens is the number that matters. It is your genuinely retained base, and it is the only defensible input to a lifetime value calculation. Using an early-period churn rate in an LTV formula produces a number that is far too pessimistic; using an uncapped late-period rate produces one that is absurdly optimistic.
Compare cohorts at the same age, always
The single most useful read is a vertical one: take month 6, and look at it across every cohort that has reached month 6. If newer cohorts retain better at the same age, something you did worked. If they retain worse, something broke, and you can usually name the quarter.
Retention above 100% is expansion, and it is good
Revenue-based cohorts can exceed 100% because existing customers grow. A cohort spending more in month 12 than in month 0 has expanded more than it has churned.
This is the same effect that lets net revenue retention exceed 100%, and it is one of the strongest signals in a B2B business. Count-based cohorts cannot exceed 100%, which is one reason to build both.
Revenue cohorts or customer cohorts?
| Revenue cohorts | Customer cohorts | |
|---|---|---|
| Best for | Financial planning and LTV | Product and onboarding questions |
| Captures expansion | Yes | No |
| Can exceed 100% | Yes | No |
| Weakness | One large account can mask many small losses | Treats a $500 and a $50,000 customer identically |
Build both. They answer different questions, and the gap between them is itself informative: strong revenue retention alongside weak logo retention means you are losing small customers and growing large ones, which is a strategy question rather than a problem.
Four things only a cohort view shows
Whether growth is masking churn
Rapid new-customer acquisition holds total revenue up while the existing base erodes underneath. A revenue chart shows a healthy line. The cohort triangle shows the erosion immediately.
Which acquisition period went wrong
A single bad cohort usually traces to something specific: a pricing experiment, a channel that brought poor-fit customers, a quarter when onboarding was understaffed. Blended metrics smear that signal across every period so it never becomes actionable.
Whether onboarding changes worked
Any change to onboarding shows up as a difference in the early-period slope for cohorts acquired after the change. This is the cleanest natural experiment most companies have available.
The real shape of your LTV
The textbook LTV formula divides margin-adjusted revenue by churn rate, implicitly assuming that rate holds forever. At low churn this produces enormous, meaningless numbers. The cohort curve gives you an empirical alternative: sum the observed retention across the periods you have, and cap the extrapolation at a horizon you can defend, 36 or 60 months.
Common mistakes
- Reading down a column across ages. Newer cohorts have fewer periods. Only compare the same age.
- Drawing conclusions from thin cohorts. A cohort of four customers will look volatile regardless of what is happening. Note cohort size alongside the percentages.
- Mixing contract lengths. Annual and monthly customers have completely different retention shapes. Segment them or the blended curve describes neither.
- Ignoring seasonality in acquisition. A cohort acquired in a promotional month may be structurally different from one acquired in a normal month.
- Building it once. Cohort analysis is only valuable as a repeated measurement. A single snapshot tells you where you are and nothing about where you are going.
How many periods you need
Six periods to see a curve shape. Twelve before saying anything confident about where it flattens. With fewer than six you can still identify a badly performing cohort, but you cannot distinguish a steep early drop that stabilises from a genuine long-term decline, and those imply very different things.
Where the analysis stops
The triangle tells you which cohort underperformed. It does not tell you why. Answering that means joining the cohort to the channel it came from, the price it paid, the onboarding it received, and the support tickets it raised, which is a different dataset every time the question is asked, and usually a project rather than a query.
That is the pattern across most of financial analysis: producing the measurement is tractable, and explaining it is where the time goes.
Choosing the cohort definition first
The most common cohort mistake is not analytical: it is grouping by the wrong thing and discovering it three hours in. Decide before you build:
- Signup month is the default and right for most SaaS. It answers "is what we are selling now better than what we sold a year ago?"
- First-payment month matters where trials are long or onboarding is slow, because signup and revenue can sit quarters apart.
- Contract-start month is the honest choice for annual enterprise deals, where signup date is an artefact of the sales process.
- Acquisition channel cuts across all of the above and is usually where the real finding is: blended retention often hides one excellent channel and one that should be switched off.
Reading the triangle
Three shapes account for most of what a retention triangle tells you:
- Where the curve flattens. Steep early decay followed by a flat plateau is a healthy product with an onboarding problem. The plateau level, not the month-one drop, is what determines lifetime value.
- Whether recent cohorts sit above older ones. Read the triangle diagonally. Improving cohorts mean the product or the targeting got better; the aggregate number will lag this by quarters.
- Revenue cohorts above 100%. Entirely normal and it means expansion is outrunning churn in that cohort, but it can also mask that you are losing many small customers while growing a few large ones. Always read revenue and logo retention side by side.
Two traps that produce wrong conclusions
Immature cohorts. A cohort three months old cannot tell you anything about twelve-month retention, and including it in an average drags the average toward whatever recent performance looks like. Cut the triangle at the last fully-mature period for any headline number.
Blended churn. A single churn percentage averages a good cohort with a bad one and hides the trend that matters. Two businesses with identical blended churn (one improving every cohort, one deteriorating) need opposite decisions, and the blended number cannot distinguish them.
What to do with the finding
A retention triangle is only useful if it changes something. Three actions it commonly justifies:
- Reallocate acquisition spend toward the channel whose cohorts plateau highest, which is often not the one with the lowest CAC.
- Fix onboarding where the month-one drop is steep but the plateau is healthy: that shape says the product works and the first fortnight does not.
- Reprice or re-segment where a whole cohort decays uniformly, which usually indicates the segment was mis-sold rather than under-served.
The analysis takes an afternoon of pivot tables by hand and about a minute pasted into a tool. What it does not do by itself is tell you which of the three you are looking at: that still needs somebody who knows what changed in the business that month.
How often to rebuild it
Quarterly is right for most mid-market SaaS. Monthly adds noise without adding signal, because a single month rarely moves a mature cohort, and annual is too slow to catch a channel going bad.
The exception is after a deliberate change: new pricing, a new segment, a different acquisition channel. Those cohorts are worth watching from month one, because that is precisely the case where you want to know early whether the change worked.
Frequently asked questions
What is cohort retention analysis?
Grouping customers by the period in which they first purchased, then measuring what each group spends in every subsequent period as a percentage of their first-period spend. It shows whether retention is improving across successive cohorts, which a blended churn rate cannot.
How is cohort retention different from churn rate?
Churn is a single figure for a period, blending every customer regardless of when they joined. Cohort retention tracks each joining group separately over time. Churn tells you what happened; cohorts tell you which direction you are heading and which acquisition period went wrong.
Should cohorts be built on revenue or customer count?
Revenue for financial planning, because it captures expansion and contraction as well as outright loss. Customer-count cohorts are better for product and onboarding questions. Revenue cohorts can exceed 100% retention; count-based cohorts cannot.
How many periods do you need for cohort analysis to be useful?
At least six to see a curve shape, and twelve before saying anything confident about where it flattens. With fewer than six you can still identify a badly performing cohort, but you cannot yet distinguish a steep early drop from a genuine long-term decline.