Are our predicted lifetime value models accurate against actual customer behavior?
№ 065 · pLTV vs realized LTV
Definition
MetricPredicted lifetime value against realised lifetime value for the same cohort at the same horizon
Unitpredicted / actual ratio, plus rank correlation between predicted and actual
Two lines on the same chart: predicted lifetime value (from a model at signup) vs actual LTV over time. Watch the gap. Build by saving model prediction at signup and comparing to actual gross profit accumulated. Example: pLTV $120; realized at M12 $96 = 20% optimism (recalibrate model).
- LTV
- Lifetime value. Average gross profit per customer per period divided by that period churn rate, or summed over expected life.
Benchmarks
| Bottom 30% | Median | Top 30% |
|---|---|---|
| predicted off by 2x to 3x; rank correlation below 0.5, meaning the model does not usefully order customers | predicted within roughly 20% to 30% of actual at 12 months; rank correlation 0.5 to 0.7 | predicted within roughly 10% of actual at 12 months; Spearman rank correlation above 0.7 |
DTC and ecommerce predictive LTV models scored against realised cohort outcomes, 2024 to 2026. Definition used: compare day-30 predictions against realised value at day 365 for the same cohort, and score both calibration (level) and rank ordering (discrimination). · Confidence is low and should stay low. No Tier 1 dataset publishes a distribution of pLTV accuracy across brands, because accuracy is a property of each brand's own model and data rather than of an industry. The bands here are assembled from published methodology guidance rather than from a measured sample, and the two commonly quoted accuracy figures (roughly 60% by day 7 and above 80% by day 30) come from vendor and practitioner sources without disclosed test methodology, so treat them as claims rather than benchmarks. The one well-sourced anchor is directional: McKinsey research has repeatedly found company-level LTV estimates off by a factor of 2 to 3 against realised cohort performance in the first 24 months. The practical guidance across sources is consistent even where the numbers are not: rank ordering matters more than point accuracy for acquisition bidding, and a model must be retrained at least quarterly because product mix, pricing and competition shift the underlying behaviour.
Category split omitted: Model accuracy is a property of each brand's data and modelling choices, not of a vertical, and no independent source publishes accuracy distributions by category.
When it looks bad
The scatter of predicted against realised sits consistently above the 45-degree line, and the gap widens for the highest predicted deciles, which means the model is most wrong exactly where it is being used to justify the highest acquisition bids.
Day-30 predictions for the top decile average $410 against realised 12-month value of $214, while the bottom decile is close to accurate. CAC ceilings were set off the top-decile prediction, so the brand overpaid on its highest-volume lookalike audiences for three quarters.
What to do about it
- Score the model on rank correlation as well as level. Point accuracy matters less than correct ordering for bidding decisions, and rank correlation above 0.5 is usable while above 0.7 is strong; a model can be badly calibrated and still allocate budget correctly.
- Run a fixed back-test each quarter: freeze day-30 predictions for a cohort, then compare against realised value at day 180 and day 365 for that same cohort. Without this the model degrades silently as mix and pricing shift.
- Use predicted value bands rather than point estimates when setting CAC ceilings, and design decisions that survive being wrong by 30%. McKinsey-cited evidence puts company-level LTV estimates 2x to 3x off realised performance in the first 24 months.
- Retrain at least quarterly and alert on drift. A model trained on a pre-change discount regime or product mix will keep predicting the old behaviour long after the cohorts have changed.
Sources
- Exactius Spearman rank correlation between predicted and actual LTV above 0.5 indicates useful predictive power and above 0.7 is strong; validation should compare day-30 predictions against actual LTV at day 180 and day 365 for historical cohorts exacti.us ↗
- AdLibrary (citing McKinsey and Bain) McKinsey research has repeatedly shown company-level LTV estimates are off by a factor of 2 to 3 against realised cohort performance, particularly in the first 24 months; changing any LTV input by 10% moves the output 25% to 40% adlibrary.com ↗
- Finsi Recommends pLTV ranges rather than hard cutoffs, quarterly retraining and continuous accuracy monitoring with drift alerting, because a model trained on 2024 data will not perform on 2026 customers as mix, pricing and expectations shift finsi.ai ↗
- Decile Accuracy is also timeliness: a model predicting 12-month LTV with 85% accuracy but only after 9 months of data is nearly useless for acquisition bidding; buyers should ask to see how predictions compare to actuals on test cohorts decile.com ↗
- CLIMB A well-calibrated predictive model reaches roughly 60% accuracy by day 7 after first purchase and above 80% by day 30, using BG/NBD for purchase timing and Gamma-Gamma for spend climbtheladder.com ↗
- Lifetimely (AMP) Predictive model produces month-to-month LTV projections by segment, with LTV drivers reported by product, promotion and channel for comparison against realised cohort behaviour lifetimely.io ↗
- MercuryMinds Blended LTV:CAC of 3.2:1 concealed active segments running at 1.9:1 and worsening, because aggregate LTV is a lagging average across cohorts with structurally different curves mercuryminds.com ↗