Beliefs, Not Forecasts: Making Private-Markets Decisions When the Data Can't Settle Them
A standard private-equity diligence package contains a recognizable set of statistics: a manager's IRR, TVPI, DPI, public-market equivalent, perhaps a risk-adjusted alpha and a Sharpe ratio — usually reported with confidence intervals and stated to two or three decimals. The implicit message is that the data have spoken and the manager's risk-adjusted performance has been measured.
That message is wrong in a specific way that matters for the decision in front of an allocator.
The literature disagrees with itself
The academic literature on private-equity performance, working with the entire universe of available funds, reaches conclusions that disagree by amounts comparable to the thing being estimated. One well-cited study finds buyout outperformance of about three percentage points a year. Another, using an overlapping sample, finds net-of-fees underperformance of about three points a year — six on a risk-adjusted basis. A third documents that conclusions differ systematically across the major commercial datasets.
The same asset class, examined by competent researchers using overlapping samples, produces point estimates that differ by six percentage points or more. If the literature cannot agree on the sign of buyout alpha using thousands of funds, the inference that any single GP's track record can be reliably risk-adjusted by parametric methods is not credible. This is not a problem of effort or sophistication. It is a structural property of the data.
Why the data are structurally too thin
Three features of private-market data create the disagreement, and none of them goes away with more observations:
- Smoothing. GP-reported NAVs are appraisal-based, not transactable. They lag public-market price discovery, understate volatility, and bias risk estimates toward zero. Corrections exist, but each requires the analyst to specify a model of the smoothing — and different specifications give different answers.
- Selection. Databases are populated by surviving managers who choose to report. Winners backfill their track records on entry; failures wind down without final reporting. The selection is correlated with the very quantity being estimated, so a larger sample produces a more precise estimate of a biased number — not the population number.
- Specification. Risk-adjusted performance depends on choices the data do not determine: which factor model, which peer universe, which smoothing correction, which benchmark. Each admits several defensible alternatives, and estimates under different choices typically differ by more than the reported sampling error.
The standard error printed next to a private-market performance estimate captures only the sampling component of total error. It omits the smoothing, selection, and specification components — which usually dominate. The conventional confidence interval is therefore an understatement, often a large one, of the true uncertainty.
What this means for the decision
An allocation question — does this manager clear my reservation hurdle? — needs precision substantially tighter than the literature's own internal disagreement. No procedure that squeezes a single point estimate out of a smaller subset of the same population can deliver that precision.
So the decision is made under belief, not under data-supported inference. The conventional framing hides this by attaching point estimates and confidence intervals to parameters the data do not actually identify. The honest framing exposes the role of belief — and then asks the allocator to defend that belief explicitly. That is the whole move.
The reframing: from estimating probabilities to specifying beliefs
The framework rests on a distinction the conventional approach collapses: the difference between the outcomes an LP cares about and the beliefs the LP holds about them. What outcomes matter — where the line sits between a disappointing result and a decisive one — is a property of the LP's own preferences. How likely each of those outcomes is, for a particular manager, is a matter of belief. The data inform the second; they do not determine it.
That reframing changes what the LP is being asked to produce. Instead of a single risk-adjusted number carrying a false air of measurement, the LP works from what it genuinely holds: a considered view about a manager, anchored in peer evidence and disciplined by what the LP can actually defend. Because the data do not pin that view down to one answer, the honest posture is to carry the range of views the LP regards as credible, rather than pretend to a precision the evidence cannot support.
The framework does not eliminate belief. It relocates it — from a hidden assumption buried inside an estimation routine to an explicit, articulated input the decision-maker and any later reviewer can see. How that credible range is constructed from peer evidence and manager-specific findings, and how a commitment decision is read off it, is the substance of the diligence procedure the framework defines — developed with a worked example in the practitioner companion and grounded formally in the academic paper.
The record it leaves
The payoff an institutional allocator should care about is the record. The commitment is documented as a falsifiable claim: these were the outcomes we cared about, this was the evidence we anchored on, this is the view we were willing to defend, and this is why the case held. A successor, an investment committee, or a post-hoc reviewer can revisit that record against what actually happened. Stating the belief up front and testing it before deciding is procedural prudence — the discipline experienced allocators already approximate informally, made explicit and auditable.
It is more honest, more auditable, and more decision-useful than the parametric apparatus it replaces — precisely because it stops pretending the data settled a question they cannot.