Mora Munoz Partners

All on the Line · Risk

The Coin Always Lands Heads

When the denominator is hidden, any result can look like proof. Theranos knew it. So does every hedge fund with a three-year track record.

← All essays
Theme

Probability and evidence

Published

12 May 2026

Reading time

8 minutes

Also on
SubstackLinkedInMedium
The Coin Always Lands Heads, All on the Line essay card

In John Carreyrou’s account of the Theranos fraud, there is a passage that allows you to understand how they were actually able to fool so many people for such a long time. A couple of employees were tasked with the job of retesting blood samples over and over to measure how much their results varied, and calculating each blood test’s coefficient of variation (CV). A test is generally considered precise if its CV is less than 10 percent. Theranos’s devices were proving unreliable, and these employees discovered that data runs failing to meet the required CV thresholds were simply discarded. The experiments were repeated until a satisfactory result appeared, and that result alone was recorded, reported, and eventually presented to investors, regulators, and patients. Carreyrou frames this as scientific misconduct, which it was — but he also translates it into mathematical terms that expose something deeper: a systematic exploitation of how probability actually works, dressed up as laboratory science.

The structure of what Theranos did can be stated precisely. If you flip a fair coin ten times, the probability that all ten land heads is exactly 1 in 1,024 — roughly 0.1 percent.1 That outcome is rare, but rare is not the same as impossible, and crucially, rare does not mean unrepeatable. If you run that ten-flip experiment enough times, the all-heads result will eventually appear. In 1,000 attempts, the probability of seeing at least one perfect run of heads is approximately 62 percent — better than even odds.2 What Theranos did was run the coin-flip experiment hundreds of times, wait for the inevitable all-heads sequence, and then present that sequence as proof that their coin always lands heads. The other 999 runs were not mentioned because they were never recorded. The denominator — the number of attempts it took to produce the result — was the one piece of information that would have made the whole picture legible, and blown the lid off, yet it was the one piece that was never disclosed.

This is not only a story about fraud. It is a story about how probability is used selectively — in science, in finance, and in the spaces between — to present results that cannot be evaluated without information that is almost never disclosed. The Theranos mechanism operates in legitimate finance with enough frequency that it deserves a structural framing. The coin flip makes that structure visible.

Two Ways to Lie With a Coin Flip

The Theranos mechanism contains two distinct mathematical sins, and it is worth separating them because each one operates independently and each one appears, in a different form, throughout financial markets.

The first sin is the hidden denominator. Theranos did not get lucky on the first attempt and mistake luck for truth. They ran the experiment many times — the passage makes this explicit — and kept running it until the favorable result appeared. That is precisely how probability works: given enough attempts, a 1-in-1,024 event will eventually occur, and in 1,000 attempts the probability of seeing it at least once is already 62 percent. The all-heads run was not a miracle. It was an inevitable consequence of sampling a distribution until its tail appeared. The fraud was in then presenting that tail event without its context — stripping away the hundreds of failed runs and offering the single success as if it had emerged from a single, unbiased trial. The denominator, the number of attempts required to produce the result, was the one piece of information that would have converted the result from compelling to meaningless. It was never disclosed because disclosing it would have ended the story immediately.

The second sin is more subtle and in some ways more damaging, because it is the error that persists even when people understand the first one. Theranos did not merely claim that their technology had worked once. They claimed it always worked — that the all-heads result revealed a permanent property of the system. This is a claim that no finite sequence of coin flips can ever support, regardless of how that sequence was generated. A run of ten heads, even if it were genuinely the first and only attempt with no selection involved, tells you nothing about what the eleventh flip will produce. The coin has no memory. A single realised outcome, however dramatic, cannot establish the distribution that generated it.

What Theranos presented as a stable, repeatable diagnostic capability was a single draw from a process they had deliberately sampled until a favorable draw appeared — and then they generalised from that one draw to a universal law about their technology. The probability of the observed result, given their methodology, was close to certain. The probability of the result they claimed — that the system reliably performed at that level — was precisely what had never been measured.

These two errors compound. The hidden denominator makes the result look like rare evidence. The false generalisation then treats that apparent evidence as a settled conclusion. Together they construct a narrative of validated technology from what is, mathematically, a single selected observation with no inferential content whatsoever.

How Markets Produce Winners Without Producing Skill

The Theranos structure — one actor, repeated attempts, hidden denominator — requires at least some degree of intention. Someone has to decide to discard the failed runs. The more pervasive version of the same mathematical problem requires no intention at all, because the selection happens automatically across a large population rather than deliberately within a single organisation.

The mechanism is straightforward. At any given moment there are thousands of fund managers each running their own ten-flip experiment — each operating a strategy with its own risk profile, leverage, and market exposure. In a population of 1,024 managers all running genuinely fair coins, the mathematics guarantees that roughly one of them will produce all-heads in any given ten-flip period. That manager did not cheat. They did not repeat their experiment until a favorable result appeared. They simply drew the outcome that a population of that size will always contain. But the investor observing only that manager’s record, without visibility into the other 1,023 records, cannot distinguish a lucky draw from genuine skill. Only the denominator — the full population that generated it — would tell them which one they are looking at, and that is precisely what is never disclosed.

This is structurally different from the Theranos case in one critical respect. Theranos required active concealment — someone had to decide not to record the failed runs. Survivorship bias in fund management requires no such decision. The failed funds simply close, their records become inaccessible, and the population that remains visible is automatically selected for strong performance. The denominator disappears not through fraud but through the ordinary mechanics of capital allocation and fund closure. The mathematical consequence is identical. An investor evaluating the visible population of funds with strong three-year records is in the same position as an investor shown only Theranos’s successful diagnostic runs — they are looking at a sample that has been selected from a much larger distribution, and the selection process has been designed, by structure rather than by intent, to show them only the tail.

The numbers make this concrete, and they are worth tracing across time because the story is more instructive than a single snapshot. There are approximately 3,000 hedge funds reporting returns at any given time, but that population is already the survivor — funds that closed are no longer in it. Over the decade from 2011 to 2020, the average hedge fund returned 5% annually against 14.4% for the S&P 500, a gap so large that $100,000 invested in the average fund grew to $160,000 while the same amount in an index fund grew to $364,000. The industry’s defenders argued that the environment changed after 2022 — rising interest rates and higher volatility were precisely the conditions their strategies needed. In 2023, the average hedge fund returned 4.4%, below even the decade average that had already been condemned as inadequate, while the S&P 500 returned 26.3%. In 2024 returns improved to 11.9% — and the S&P 500 returned 25%.3 Across the full period from 2011 to 2024, the average hedge fund has not beaten a passive index fund in a single year on a net-of-fee basis, through bull markets, bear markets, and rate cycles alike. The funds that produced the all-heads records within that period raised the capital. The funds that did not closed quietly. The investor who allocated to the visible winners was not observing skill. They were observing the tail of a distribution that the structure of the industry was always going to produce, and paying substantially for the privilege.

The within-firm version of this problem — backtesting — sits between the two cases in terms of intention. A quantitative fund testing hundreds of parameter combinations across historical data and then presenting the best-performing set is running the Theranos experiment deliberately, but without necessarily understanding it as selection. The strategy that survived the testing process looks, in the presentation, like the strategy that was identified through rigorous research. What it actually is, mathematically, is the all-heads run drawn from a large sample of attempts — the one that the process was always going to produce, because any process that searches a distribution until a favorable result appears will find one. The number of parameter combinations that failed to produce a strong backtest is not in the pitch deck because it was never written down, and because writing it down would immediately raise the question the denominator always raises: given how many attempts this required, what does the result actually tell us?

There is a more precise way to state how uninformative a short track record actually is, and the result is more extreme than you would expect. To reach conventional statistical confidence that a fund manager’s outperformance is real rather than lucky, assuming a realistic level of skill and typical return volatility, you need decades of continuous data.4 At three years you have noise. At ten years you have slightly less noise. The industry sells three-year records because that is what the capital allocation process rewards. The mathematics says a three-year record is not evidence of skill. It is evidence that the manager has been flipping long enough for a run to appear.

Proof of Returns Can Be Indistinguishable From a Coin Flip

There is a single question that, if applied consistently, would make a sophisticated observer dramatically harder to deceive across laboratory science, financial markets, and business narratives alike. That question is not “what did the data show?” — it is “how many times did you flip before showing me this?”

Markets run the same procedure with less intention and more systemic force. There is no single decision-maker at the center of survivorship bias who chooses which funds to hide. The mechanism operates through a distributed process of capital allocation and fund closure that, in aggregate, produces a population of visible managers whose records look exactly like what you would expect if you had selected them specifically for strong performance. The population that generates this effect is not conspiring. It is doing what all populations of coin-flippers do — producing some runs of heads, some runs of tails, and a small number of long consecutive streaks that, once they appear and are presented without their context, look like something other than what they are.

The denominator is not a technical footnote. It is the number that converts a result from evidence into noise, or from noise into evidence, depending on its size. Any performance record, any clinical result, any backtest, and any diagnostic threshold presented without its denominator is uninterpretable — a result whose meaning cannot be established until you know how many attempts it took to produce it. The practice of omitting it is so widespread, and so rewarded, that demanding it consistently marks you immediately as someone who understands what the data actually contains. In most rooms, that remains a minority position.

— Carlos E. Mora

I wake up, I build, I repeat. No guarantees.

I work like it’s all on the line, because it is.

Family is the only true legacy.

Your name is your currency, and it must be earned daily.

Notes

1.Each flip of a fair coin has a 1/2 probability of landing heads. Ten independent flips multiply ten times: 1/2 × 1/2 × … × 1/2 = (1/2)¹⁰ = 1/1,024 ≈ 0.098%.

2.The probability of never seeing an all-heads result across 1,000 independent attempts is (1,023/1,024)¹⁰⁰⁰ ≈ e⁻⁰·⁹⁷⁷ ≈ 0.376, so the probability of seeing it at least once is 1 − 0.376 ≈ 62.4%.

3.Return data for hedge funds draws on the Barclay Hedge Fund Index and HFRI Fund Weighted Composite Index, as compiled by the American Enterprise Institute (2011–2020 figures), HFR (2023 figures), and Hedge Fund Alpha (2024 figures). S&P 500 annual return figures are total returns including reinvested dividends, sourced from standard index data for each calendar year. The $160,000 and $364,000 terminal values assume $100,000 invested at the start of 2011 compounded at the respective average annual rates through end of 2020.

4.To test whether a fund manager generates genuine alpha rather than luck, we use a t-statistic of the formt = (mean annual excess return ÷ volatility) × √years. Assuming an alpha of 2% annually and return volatility of 15%, the calculation to reach the conventional significance threshold of t = 2 is: (2% ÷ 15%) × √years = 2, giving √years = 15, and therefore years = 225. At three years the t-statistic is 0.23. At ten years it is 0.73. Neither approaches significance. If assumed alpha is raised to 5% — already a generous assumption — the required period falls to 36 years. To reach 19 years, which is the figure sometimes cited in academic literature, requires assuming alpha of roughly 6.5%, which most serious researchers would consider optimistic. The direction of the conclusion is robust across all realistic assumptions: the data horizon required to distinguish skill from luck in fund management is measured in decades, not years.

The practice

The arithmetic in these essays is the arithmetic the practice runs on a mandate.

Discuss a mandate →