Worked example: why did Q3 churn spike?
The mechanics are covered elsewhere. This walks through an actual investigation, including the part where the first answer is wrong — which is the part checkpoints exist for.
The question
Churn looked worse in Q3 than Q2 and nobody knows why. You have data/customers.csv, one row per account, with status, region, plan and a signup cohort.
The plan: get to a clean base table, then test hypotheses against it. The base table is the thing worth protecting — it is the expensive part, and every hypothesis will start from it.
Step 1 — build the base
let raw = load.csv("data/customers.csv")
let active = filter(raw, status, "active")
let report = select(active, [id, name, region, plan, cohort])Three steps, three checkpoints. Watch the row count in the inspector as you go — raw is 12,480 rows and active is 4,210. That drop is expected here, but it is exactly the kind of number worth noticing: if a filter removes far more than you thought, you have found a data problem rather than an answer.
Name the last one deliberately. report is the checkpoint every branch below forks from, and its name is what you will see in the graph.
Step 2 — first hypothesis: it is a pricing problem
The obvious guess is that the Q3 price change drove people away. Test it by branching from report rather than editing it — that keeps the base intact whatever happens.
let priced = apply(report, pricing_test)
export.csv(priced, "data/pricing-test.csv")The result does not support it. Churn is flat across plans, and the accounts that left were not disproportionately on the new price.
This is the moment that matters. In a notebook you would now edit those cells away and lose the evidence that the hypothesis was tested. Here you rewind to report and the branch stays where it is, labelled, with its results attached. If someone asks in three weeks whether it was pricing, the answer is on screen.
Step 3 — rewind and try again
Select the report checkpoint. State is restored — the 12,480-row load does not run again, which is the whole point. Now branch a second time on a different idea: a cohort effect.
let windowed = window(report, days, 90)
let cohort = filter(windowed, cohort, "q3_signups")
export.csv(cohort, "data/cohort-test.csv")This one moves. Accounts that signed up in Q3 retain 12% better over 90 days than the overall base — meaning the spike is not new customers leaving. Something is happening to older accounts.
Step 4 — narrow it
Branch again, this time from the cohort step rather than the base, and split by region:
let byRegion = group(cohort, region)
export.csv(byRegion, "data/region-split.csv")EU accounts are up 4% while the rest are flat. Now you have something specific enough to act on, and a graph showing exactly how you got there.
What the history looks like at the end
Four branches off one base:
- main — load, filter, select, export
- pricing_test — tested, discarded, still on record
- cohort_test — +12% retention
- region_split — +4% in EU
Nothing was overwritten and nothing was recomputed twice. The discarded branch is as much a part of the answer as the one that worked — it is the reason you can say pricing was ruled out rather than never checked.
The habits worth taking from this
- Get to a clean base first, and name it. Everything forks from it, so a good name pays for itself in the graph.
- Branch per hypothesis, not per edit. One idea, one branch, one exported result.
- Watch row counts between checkpoints. The inspector comparison catches bad filters long before the results table does.
- Keep the branches that failed. They cost nothing and they answer “did you check…?”.
Next: the statements used here are in the VeraScript reference, and the range and branching mechanics are in Checkpoints.