All on the Line · Probability
Why does professional tennis produce new records every week? The same combinatorics explains the explosion of ETFs, with very different consequences.
Combinatorics
26 May 2026
14 minutes

On Sunday, May 17, 2026, Jannik Sinner — the number one tennis player in the world — beat Casper Ruud 6-4, 6-4 at the Foro Italico to win the Italian Open. And an impressive headline record was announced: six consecutive ATP Masters 1000 titles, a streak dating back to the Paris Masters in November of last year. A feat never achieved before in the sport of tennis. But inside that single headline, nested like Russian dolls, were eight more records broken or equaled in the same stretch. Sinner had extended his consecutive Masters 1000 match-win streak to 34, surpassing Djokovic’s previous mark of 31. He had completed the Career Golden Masters — all nine distinct Masters 1000 titles won across a career, joining only Djokovic in the sport’s history. He had swept the three clay Masters of a single season — Monte-Carlo, Madrid, Rome — the first to do it since Nadal in 2010. He had become the first Italian to win the Italian Open since Adriano Panatta in 1976, a fifty-year drought. He had completed the Sunshine Double without conceding a set in either tournament, a first in tennis history. He had been the first player ever to win five consecutive Masters 1000 events, then extended that record to six in Rome itself. He had broken Djokovic’s record for most consecutive sets won at Masters 1000 events, reaching 37 before the streak ended. He had surpassed Federer’s career match win rate along the way.1
And those are just the ones that made the headlines — beneath them sat a longer list of marks reached and equaled, from passing Alcaraz in the Big Titles standings to joining the eight-man club of Sunshine Double winners to top-tier rankings on serve and return leaderboards that broadcasters track separately from the headline records. And in the next few days after the Rome final, I kept stumbling upon more records: that Sinner has held the world number one ranking for 72 weeks at age 24, against Djokovic’s 26 weeks at the same age; that he is the first ever to win the first four Masters 1000 events of a single season. Numerous named records, an unknown longer tail, one streak, and Roland Garros is on the way. It started on May 24, and by the time you read this, there will probably be more records added. The supply of new records about familiar players appears to be infinite, and the reason is worth understanding.
Tennis is unusually transparent. Every match is broadcast, every shot is tracked, every statistic is logged in real time, the result is visible to anyone watching. A player wins or loses in front of everyone, and what happens between the first serve and the final point is recorded with a granularity that would be unimaginable in most other domains of human performance. This is why tennis is a useful place to study statistics — because there is no underlying performance the audience cannot see for itself. The records sit on top of a foundation the audience has already inspected. What broadcasters and statisticians do with that foundation is to slice it — to pick a combination of conditions (a surface, a tournament tier, a round, a streak, a statistic) under which some player stands alone. Every record you have ever heard in tennis is the result of a slicing. Underneath that foundation sits a dataset. Jeff Sackmann, who runs Tennis Abstract and the Match Charting Project, maintains the most widely-used historical record of professional men’s tennis on a public GitHub repository.2 It is the dataset that serious tennis analysts, journalists, and academic researchers actually use, and it has a precise structure worth examining.
Every match in the database is described by forty-nine columns. There are three recorded surfaces — clay, grass, and hard — though the sport itself plays on four, since indoor hard and outdoor hard behave differently enough that commentators treat them as separate. There are seven recognized tournament tiers — Grand Slam, ATP Finals, Masters 1000, ATP 500, ATP 250, Davis Cup, Olympics — though the dataset compresses 500s and 250s into a single category. There are seven main-draw rounds, from the first round of 128 through the final, plus round-robin play at the ATP Finals and the bronze medal match at the Olympics. There are two match formats, best of three and best of five sets. There are eighty-two distinct nationalities appearing in a single recent year. There are three handedness matchups — righty against righty, righty against lefty, lefty against lefty — which matter because the spin patterns differ enough that coaches plan against them. There are five common ranking tiers — top-5, top-10, top-20, top-50, outside top-50 — which is how the sport itself talks about who is playing whom. There are five age brackets. And then there are the in-match statistics — aces, double faults, serve points, first serves in, first-serve points won, second-serve points won, service games, break points saved, break points faced. Nine statistics, each of which can be conditioned on a specific numerical threshold of the kind broadcasters actually use, like “more than 20 aces” or “above 70% on first-serve points won.”
Multiply the categorical dimensions and the number grows quickly. Four surfaces times seven tournament tiers times seven main-draw rounds times two match formats gives 392 distinct match contexts before you have conditioned on anything about the players themselves. Add five age brackets, five ranking tiers, three handedness matchups, and the number reaches roughly 29,000. Now add the in-match statistical conditions. Each of the nine statistics tracked per match gives perhaps four meaningful round-number thresholds — the ten-ace mark, the fifteen-ace mark, the twenty-ace mark, the twenty-five-ace mark, and so on for each statistic — which gives 36 single-statistic slicings per context. Multiplied through, the realistic category space sits at approximately one million distinct, intelligible, defensible slicings of the historical record of professional tennis. And this is conservative — it counts only the dimensions documented in the standard academic dataset, treats single-match statistics in isolation rather than in combination, and excludes the structural conditions broadcasters use constantly (consecutive-win streaks, head-to-head structure, career aggregates, same-tournament history) which push the count substantially higher.
That is the supply side. Now the consumption side. The ATP plays roughly three thousand main-draw tour-level matches per year. The slicing space is one million. Even at this conservative count, the supply of record categories exceeds the annual match flow by more than three hundred to one. The year’s matches, taken together, can fill at most a tiny fraction of the available space. And a single match is not a single point in this space — it occupies dozens of slicings at once, because it is on a specific surface, in a specific tier, in a specific round, between players of specific handedness, with specific statistical outputs. The chance that every one of those dozens of slicings has already been filled by a prior match is essentially zero. So nearly every match, by structural necessity, produces at least one slicing in which it stands alone. The asymmetry is structural and it is widening. The match count grows slowly — the tour is at or near its calendar limit, and the addition of a new tournament is rare and politically contested — while the dimension count grows faster, as the ATP and its broadcast partners continually add new measurement axes to what professional tennis tracks. Performance Rating, Shot Quality, Steal, Conversion — each one is a new dimension folded into the existing space, multiplying the count of intelligible slicings. The gap between supply and demand is not closing. It is opening.3
This is what is happening when you watch a broadcaster announce, twice a match, that some player has just achieved something nobody has ever achieved before. The records are not being invented. They are being found in a space large enough that finding is always possible. The combinatorial mathematics of the dataset guarantees that every significant match produces multiple unprecedented intersections, and the broadcaster’s job is to surface the most resonant of them in time for the next changeover. There is no shortage. There has never been a shortage. There cannot be a shortage. The supply is the structure. And the math does not invent the greatness. It articulates a greatness the player has earned, across every dimension the dataset can see.
Alejandro Davidovich Fokina is a Spanish tennis player with a career-high ranking of 14, reached in late 2025, who has never won an ATP title. He turned professional in 2017, and across 321 main-draw tour-level matches between 2019 and 2026 he has reached five finals and lost all five. He is recognizable to anyone who follows tennis closely — a baseliner with an aggressive game and a fluent dropshot — and almost invisible to anyone who does not. His records do not circulate the way Sinner’s do. And yet, when you run the same combinatorial machinery against his match data that the official statisticians run against Sinner’s, the records appear immediately. Three of them are worth naming.4
The first sits in the same broadcaster shape used to describe Sinner. In 2025, Davidovich Fokina reached the final of four ATP 500 or 250 hard-court tournaments — Delray Beach in February, Acapulco later that month, Washington in July, Basel in October. Across the eight tour seasons documented in Sackmann’s dataset, only nine players have done this in a single year. The other eight are Sinner (twice, in 2021 and 2023), Medvedev (twice, in 2019 and 2023), Auger-Aliassime (twice, in 2022 and 2025), Rublev, and de Minaur. The grammar of the record is identical to the grammar broadcasters use for Sinner. The filter stack is short: hard court, ATP 500/250 tier, finals reached, single calendar year. The slicing produces a club of nine players in eight years, and Davidovich Fokina is one of them.
The second is sharper because it conditions on nationality. Davidovich Fokina is the only Spanish player in the dataset’s coverage to reach four hard-court ATP 500 or 250 finals in a single season. Spanish tennis is overwhelmingly defined by clay — Nadal, Alcaraz, Bautista Agut, Carreño Busta — and the historical pattern of Spanish hard-court performance reflects that emphasis. A Spanish player reaching this density of hard-court finals is structurally unusual in the sport’s national-tradition terms, and the slicing produces him as the unique holder of the record within his country during this period. This is precisely the kind of nationality-conditioned record broadcasters surface constantly when an elite player satisfies it — first Italian since, first American since, first Australian to — and the essay’s introduction leaned on three records of this exact shape about Sinner. The shape works for Davidovich Fokina too.
The third is the kind of single-match record that broadcasters love most. On April 11, 2022, in the third round of the Monte Carlo Masters, Davidovich Fokina beat Novak Djokovic 6-3, 6-7(5), 6-1. In that match he broke Djokovic’s serve nine times. Across Djokovic’s entire professional career from 2005 to 2026 — 892 best-of-three matches in which the standard statistical record exists — nobody has ever broken Djokovic’s serve more times in a best-of-three match than Davidovich Fokina did that day in Monte Carlo. The next highest is eight, by Musetti at the same tournament a year later. After that, seven, by Nadal in 2009. Davidovich Fokina holds the all-time record for breaking the serve of the most accomplished server of his era, in the standard professional match format, across the entirety of that server’s career. It is a record that most people will never hear about, but nonetheless it lives in the data.
What is true of Davidovich Fokina is true of every player who has played at this level for any length of time. The slicing space is dense enough that the act of pointing the machinery at any reasonably accomplished player produces real records of the same broadcaster-grammar shape that surfaces routinely about the elite. The records about Sinner are not a different kind of fact from the records about Davidovich Fokina. They are the same kind of fact, surfaced by the same kind of math. The math is democratic. The broadcasting is not.
What happens in tennis happens in finance, at much larger scale and over a much shorter history. Over the past two decades the financial industry has built the same kind of high-dimensional category space the tennis world built across a century — and it has built it faster, more deliberately, and with more economic momentum behind it. Stock indices — the named rankings of stocks like the S&P 500 or the Nasdaq 100, used as benchmarks against which funds and investors measure performance — have proliferated at extraordinary scale. The Index Industry Association, the trade body that tracks them, reports that there are now roughly 3.3 million stock indices in existence globally. The World Bank counts roughly 43,000 publicly listed companies in the world.5 The ratio is more than seventy indices for every listed company.
The growth is concentrated in the products built on top of those indices. Exchange-traded funds — or ETFs — are, in their most common form, investable wrappers that hold the stocks in a particular index in roughly the same proportions, so that buying a share of the fund gives the investor exposure to the whole index in one transaction. Other ETFs use derivatives or active management to pursue strategies that do not simply mirror an index. As of the end of April 2026, there were 16,605 ETFs globally, with 32,401 listings, from 1,004 providers across 86 exchanges in 66 countries. In 2025 alone, 2,759 new ETFs launched globally, roughly 230 every month for a year.6 The category space of investment vehicles is expanding faster than the underlying universe of companies they invest in. The same single stock now sits inside a startling number of indices simultaneously. Apple is a component of the S&P 500, the S&P 100, the Nasdaq 100, the Nasdaq Composite, the Russell 1000, the Russell 3000, the S&P 500 Information Technology, the S&P 500 Growth, the S&P 500 Value, the S&P 500 ESG, the MSCI World, the MSCI ACWI, dozens of MSCI sector and factor and thematic indices, FTSE’s global series, and the hundreds of custom indices built by ETF providers around themes like artificial intelligence, megacap tech, dividend growth, and the rest. By construction, Apple is a top-weighted component of some index, in some time window, every day of the year.
The dimensions the industry slices on have grown unusually narrow. Among the funds registered with the SEC by a single mid-sized provider, Roundhill Investments, are the Roundhill GLP-1 & Weight Loss ETF, the Roundhill Daily 2X Long Magnificent Seven ETF, the Roundhill Uranium ETF, the Roundhill Meme Stock Covered Call ETF, the Roundhill Humanoid Robotics ETF, and a series of single-stock weekly-pay derivative ETFs covering Apple, Tesla, Nvidia, MicroStrategy, Reddit, and a dozen others.7 Each of these is a category that did not exist five years ago. Each is now an investable wrapper around a slicing of the underlying universe that someone constructed deliberately.
This deliberate construction of categories has, in one specific industry practice, been taken to its mathematical limit. It is called self-indexing. Nearly twenty percent of US ETFs track an index created by the same firm that manages the fund. The provider designs the index, defines the methodology, sets the rebalancing schedule, and chooses the universe, and then launches a fund that tracks the index they themselves built. Rick Redding, founding CEO of the Index Industry Association, has asked whether such indices have any economic intuition behind them other than back-tested data.⁸
This deliberate construction of categories has, in one specific industry practice, been taken to its mathematical limit. It is called self-indexing. Nearly twenty percent of US ETFs track an index created by the same firm that manages the fund. The provider designs the index, defines the methodology, sets the rebalancing schedule, and chooses the universe, and then launches a fund that tracks the index they themselves built. Rick Redding, founding CEO of the Index Industry Association, has asked whether such indices have any economic intuition behind them other than back-tested data.8
In tennis the records are the end. A fan reads that Davidovich Fokina has broken Djokovic’s serve more times in a single best-of-three match than anyone else, and the fan enjoys the record. The record is the object of consumption. It does not need to be the basis of any further decision. The fan does not allocate capital based on it. The fan does not change behavior in any way the record was supposed to inform. The record is entertainment, and entertainment does not require a substrate of evaluation. The same fan can read records about Sinner and records about Davidovich Fokina and take pleasure in both, even knowing that one player is much greater than the other, because the records do not need to discriminate between the players to do their work. They only need to be true.
In finance the records are a means to a separate end. The investor who reads that some ETF is top-quartile in mid-cap value over five years is not consuming the record for its own sake. They are using the record as input into a decision about where to allocate capital that will compound over decades. And here the underlying performance behind the slicing matters way more than it does in tennis. If Davidovich Fokina and Sinner were ETFs and you saw their respective records — Davidovich Fokina’s all-time Djokovic-break record, Sinner’s six-Masters streak — and you allocated your retirement savings to Davidovich Fokina because his record sounded compelling, you would do meaningfully worse over a thirty-year horizon than the investor who allocated to Sinner. The records do not tell you that. The records by themselves do not distinguish between the two players in the way that matters for capital allocation. Only the underlying performance does, and the underlying performance is one step further away than the record.
This is why the work in finance does not stop at the surfaced record. The investor who reads the marketing material has done a fraction of the work the situation requires. The rest of the work — examining the underlying universe of alternatives the slicing was drawn from, comparing the fund against benchmarks the provider did not choose, weighing the costs of the wrapper against the substance of the strategy — sits behind the surface and has to be reached for. In tennis no equivalent work is required because no equivalent stakes are involved. In finance the stakes are the audience’s own future, and the records by themselves are not enough to navigate them.
The mathematics of slicing is everywhere. What changes between domains is what the audience is trying to do with the slicing’s output. In tennis the records are the point and the audience can take them as they come, because the point of the records is the records. In finance the records are not the point. They are a starting place for work the audience still has to do. A 401(k) statement shows the funds that won this year. The list is real. The records are real. And it tells you almost nothing about which fund will compound your savings between now and the year you retire. The cost of stopping at the statement is borne by the audience alone.
— Carlos E. Mora
I wake up, I build, I repeat. No guarantees.
I work like it’s all on the line, because it is.
Family is the only true legacy.
Your name is your currency, and it must be earned daily.
1.All record claims about Sinner’s 2026 Masters 1000 streak are sourced from the ATP Tour’s official coverage at atptour.com (specifically the Rome final and Career Golden Masters reports), with secondary aggregation from the “2026 Jannik Sinner tennis season” article and the historical comparisons from “ATP Tour records” on Wikipedia.
2.Jeff Sackmann’s tennis match dataset is available at github.com/JeffSackmann/tennis_atp, published under a Creative Commons license with attribution. The 49-column schema and the dimension values cited in this section come from the 2024 match file in that repository, which is the most recent complete season in which all tournament tiers appear, including the Olympics. Sackmann’s repository is the standard data source used by Tennis Abstract, FiveThirtyEight’s tennis coverage, and most quantitative tennis writing.
3.The combinatorial estimate is defensible at each step. The categorical context dimensions — surface, tournament tier, round, format, age bracket, ranking tier, handedness matchup — are all documented in the dataset with their value sets exhaustively listed, giving 4 × 7 × 7 × 2 × 5 × 5 × 3 ≈ 29,000 contextualized match types by direct multiplication. The in-match threshold count of four per statistic is an estimate, not a count, but is grounded in observable broadcaster practice: round-number thresholds at meaningful intervals (the ten-ace mark, the twenty-ace mark, the seventy-percent first-serve mark) are what broadcasters actually cite, and four such thresholds per statistic is a reasonable middle estimate between the two or three that appear in routine commentary and the five or six that appear across a full year of records. Nine statistics times four thresholds gives 36 single-statistic slicings per context, and 29,000 × 36 ≈ one million. The structural conditions beyond the single match — streaks, head-to-head, career aggregates, same-tournament history — are real and used constantly but interact with the categorical dimensions in ways that resist clean combinatorial estimation, so they are excluded from the headline number. The forward-looking ratio of more than three hundred to one is the slicing space divided by the approximate annual tour-level match count of three thousand. The figure for new dimensions added by the ATP (Performance Rating, Shot Quality, Steal, Conversion) refers to the Tennis Data Innovations leaderboard, accessible through atptour.com/en/stats.
4.All three records about Davidovich Fokina were computed directly from Jeff Sackmann’s dataset (see note 2). His career statistics — career-high ranking of 14, five career main-draw finals, 321 main-draw tour-level matches from 2019 through 2026 — are aggregated from his complete tour-level match record across those years. The 2025 hard-court finals club was identified by scanning all ATP 500 and ATP 250 hard-court finals across the dataset’s coverage from 2019 to 2026 and listing every player who reached four or more such finals in a single calendar year; nine players satisfy this filter, including Davidovich Fokina in 2025. The Spanish-only framing was verified by restricting the same query to players with Spanish nationality (winner_ioc or loser_ioc equal to ESP). The Djokovic break-point record was verified by scanning every Djokovic best-of-three match in the dataset from 2005 to 2026 — 892 matches with complete break-point statistics — and computing the number of times Djokovic was broken in each, defined as break points faced minus break points saved.
5.The Index Industry Association is the global trade association for index providers. The figure of approximately 3.3 million stock indices is from the Index Industry Association’s annual benchmark survey, published by the association and reported in outlets including Pensions & Investments. The figure of approximately 43,000 publicly listed companies globally is from the World Bank’s “Listed domestic companies, total” indicator.
6.ETF count, listings, providers, and exchanges are from ETFGI’s monthly global ETF industry report for April 2026, which placed global ETF assets at $21.91 trillion across 16,605 products and 1,004 providers. The 2025 figure of 2,759 new ETF launches is from ETFGI’s year-end 2025 reporting. ETFGI is the standard industry research firm tracking ETF flows and product counts; their reports are widely cited in financial industry coverage.
7.The Roundhill ETF Trust fund list as of late 2025 is from the firm’s SEC Form 485BPOS filings, publicly accessible via EDGAR. Roundhill is one provider among many; comparable narrow-slicing fund families are operated by Defiance, GraniteShares, YieldMax, and others.
8.The figure that nearly 20% of US ETFs track a proprietary self-indexed benchmark is from academic research published in 2024 on self-indexing prevalence in US ETF markets, with corroborating coverage in ETF industry trade publications. Rick Redding’s quote appeared in Pensions & Investments coverage of the self-indexing trend.
The arithmetic in these essays is the arithmetic the practice runs on a mandate.
Discuss a mandate →