
The numbers that contain everything
You’ve probably heard the claim that somewhere in the digits of is your phone number. Your birthday too, and, if you encode text as digits, every book ever written.
That claim is often presented as a fact about , but it isn’t one. We don’t actually know whether has the property needed to make it true.
The property is called normality. Roughly speaking, a number is normal in base 10 if its digits behave the way we’d expect random digits to behave in the long run: each digit occurs with frequency , each two-digit string with frequency , each three-digit string with frequency , and so on.
That last part matters. If every finite string occurs with the expected frequency, then every finite string occurs infinitely many times. So a normal number really would contain your phone number, your birthday, and any finite piece of text you care to encode.
What makes normal numbers interesting is the mismatch between what we know in general and what we can prove in particular. Almost every real number is normal. But for familiar constants such as , , and , normality is still unproved.
Making “normal” precise
Saying that the digits of a number “look random” is vague, so let’s make it precise.
Fix a base , and write a number in base as
where each digit belongs to .
The first thing we might ask is whether each individual digit occurs equally often.
Simply normal
A number is simply normal in base if every digit occurs with limiting frequency . In other words, for every digit ,
In decimal, that means the proportion of each digit tends to .
But this isn’t enough. A sequence could contain the right number of zeroes and ones while arranging them in a very non-random way. Knowing how often individual digits occur tells us nothing about how often blocks such as , , or occur.
That leads to the stronger definition.
Normal in base b
A number is normal in base if every block of digits occurs with the frequency we would expect from independent, uniformly chosen digits.
For a block of length ,
So in decimal, each digit has frequency , each two-digit block has frequency , each three-digit block has frequency , and so on.
This is the property behind the claim from the introduction. If a number is normal, then every finite string of digits occurs not merely once, but infinitely many times.
Multiplying by shifts its base- digits one place to the left, and taking the result mod discards whatever fell off the front – so has digit sequence . Equidistribution of that sequence on is exactly the statement that every digit block appears with its expected frequency. In short, is normal in base exactly when is equidistributed on 1.
Finally, normality depends on the base.
Absolutely normal
A number is absolutely normal if it is normal in every integer base .
Thus absolutely normal normal in base simply normal in base , and neither implication can be reversed in general.
There is one important relation between bases. If and are multiplicatively dependent – meaning
for some positive integers and – then normality in one base is equivalent to normality in the other2. So, for example, being normal in base is equivalent to being normal in base or base .
For multiplicatively independent bases, no such implication holds: there are numbers normal in one base but not in the other3.
A small technical point: we defined normality for numbers in , while clearly is not.
This does not matter. The integer part contributes only finitely many digits, and finitely many digits cannot change a limiting frequency. Asking whether is normal is therefore the same as asking whether is normal.
More generally, when talking about normality, we can ignore the integer part and look only at the digits after the radix point.
Borel’s theorem: almost all numbers are normal
In 1909, Émile Borel4 proved that almost every real number is absolutely normal – normal in every integer base at once.
Here, “almost every” has a precise meaning. The exceptions form a set of Lebesgue measure zero. Equivalently, if is chosen uniformly at random from , then is absolutely normal with probability .
What follows is a modern probabilistic proof, not Borel’s original 1909 argument, but it establishes the same theorem. The basic idea is simple: treat the digits of a randomly chosen number as independent random variables, apply the strong law of large numbers, and then use the fact that there are only countably many bases and finite digit strings to check.
The digits of a random number behave like independent rolls. Fix a base and choose uniformly from . Write
in base .
Suppose we ask for the first three decimal digits to be , , . The numbers with that prefix are exactly those in
an interval of length . Since is uniform on , the probability of landing in that interval – that is, – is exactly , the same number you’d get from three independent, uniform digit draws: .
Nothing about was special. Prescribing any digits pins to an interval of length , so
The joint probability factors exactly as it should for independent variables – knowing the earlier digits tells you nothing about the next one. The same counting argument works for digits at any positions, not just a prefix. Each fixed digit cuts the surviving set of ’s by a factor of , regardless of what the unfixed digits in between are free to do. So fixing restricts to a union of intervals of total length , giving the same factorization. Apart from the usual ambiguity at numbers with two base- expansions (like in base 10 – a countable, hence measure-zero, set of exceptions), the digits
are independent and uniformly distributed over .
Now fix one block. Let be some string of digits. We want to know how often appears in the expansion of .
There’s a small complication, though: consecutive occurrences can overlap. Take and : the block starting at position uses , while the block starting at position uses – they share two digits, so the corresponding events aren’t independent, and the law of large numbers doesn’t directly apply.
Split the starting positions into groups by residue modulo . For , that’s , then , then . Within the first group, the blocks use completely disjoint digits, so they’re independent – each equal to with probability . The strong law of large numbers says the fraction of these blocks equal to converges to almost surely, and the same holds for the other two groups. Since the three groups split the starting positions evenly, their combined frequency is still .
The same argument applies for a block of any length : split the starting positions into groups by residue modulo , apply the strong law within each group – where the blocks are disjoint and hence independent, each equal to with probability – and recombine. That gives
almost surely, where counts starting positions rather than digits used (counting only blocks that fit entirely within the first digits, via , gives the same limit). For any fixed base , block length , and block , the set of numbers for which the required frequency fails has measure zero.
There are only countably many conditions. Normality asks for this statement to hold for every finite block, and absolute normality asks for it to hold in every integer base.
But there are only countably many choices involved: countably many bases , countably many lengths , and finitely many strings of each fixed length in each base. Altogether, there are still only countably many conditions.
Each condition has a measure-zero exceptional set. A countable union of measure-zero sets still has measure zero.
Outside that union, every finite block has the expected limiting frequency in every base. In other words, almost every real number is absolutely normal.
That is Borel’s theorem.
The catch: we can’t prove it for the numbers we care about
Borel’s theorem tells us that almost every real number is normal, and in fact absolutely normal. But take the numbers that turn up naturally in mathematics – , , – and the situation is very different.
We do not know whether any of them is normal in base 10. In base 2, we cannot even prove the weaker statement that their zeroes and ones occur equally often5.
Explicit normal numbers are known – if we are allowed to build a number specifically for the purpose. The simplest is Champernowne’s constant6,
formed by writing the positive integers one after another. Champernowne proved in 1933 that this number is normal in base 10. Explicit absolutely normal numbers are also known7.
And it is easy to write down numbers that are not normal. Every rational number has an eventually periodic expansion in any integer base; a periodic expansion can only ever realize a bounded number of distinct length- blocks, while there are possible ones, so most blocks of that length never occur at all, which rules out normality. The usual Liouville constant,
is another example in base 10. Its ’s become increasingly sparse – the proportion of ’s among the first digits tends to , not – so even the individual digits fail to occur with equal limiting frequency.
The standard Liouville constant above is not normal, but being a Liouville number does not itself prevent normality: there are Liouville numbers that are normal, as well as ones that are normal in no base8. It is the particular digit structure of the constant above, not its membership in the Liouville numbers as a class, that makes it non-normal.
For , computation gives us plenty of evidence but no proof. In 2016, Peter Trueb analyzed the first 22,459,157,718,361 decimal digits of and, separately, the first 18,651,926,753,033 hexadecimal digits9, counting every block of length one, two, and three in each expansion. The observed frequencies were consistent with what normality predicts. But any finite computation can only examine a finite prefix. Normality is a statement about what happens as the number of digits tends to infinity, so no amount of checking alone can settle it.
We can compute the digits of and to enormous precision, prove that both are transcendental, and write down series, products, integrals, and rapidly converging algorithms for calculating them. None of those descriptions has yielded the digit-frequency estimates needed to prove normality.
The equidistribution formulation from earlier makes the difficulty concrete. To prove that is normal in base , we would need to show that
is equidistributed in . This isn’t a matter of computing more digits. It’s about understanding the long-term behavior of the fractional parts of for arbitrarily large . The familiar representations of – geometric and analytic – have not yielded a proof of the required equidistribution.
Why “almost every” tells us so little about one number
There is no contradiction between saying that almost every real number is normal and admitting that we cannot prove normality for .
A set of measure zero does not have to be empty, or even small in any ordinary sense. It can be infinite and highly structured. Measure tells us how large a set is inside the continuum of real numbers; it does not tell us whether a particular number belongs to it.
The same gap shows up with transcendence. Almost every real number is transcendental. The algebraic numbers are countable, so they have measure zero. But this fact alone tells us nothing about whether a specific number is transcendental.
Proving that is transcendental required Hermite’s work in 187310. Lindemann proved the transcendence of in 188211. And even today, we do not know whether a number as simple to write down as
is transcendental12. In fact even the weaker question, whether is irrational, is open.
The theorem that almost every number is transcendental does not help with that question. It says that a randomly chosen real number will avoid the exceptional set with probability . is one particular point, and a measure-theoretic statement about almost all points does not tell us where that point lies.
Borel’s theorem has the same limitation. It tells us that the non-normal numbers occupy a measure-zero part of the real line. It does not give us a test that we can apply to , , or .
“Almost every number is normal” tells us a great deal about the real numbers as a whole, and almost nothing about any particular number we happen to care about.
So, does contain your phone number?
We don’t know.
If is normal, then yes: your phone number appears somewhere in its decimal expansion, and not just once, but infinitely many times. The same would be true of your birthday, this article encoded as digits, and every other finite string.
Borel’s theorem tells us that this is what happens for almost every real number. Computations tell us that the digits of behave as normality would predict, in the tests performed so far.
The claim behind the phone-number story – that contains every finite sequence of digits – is still unproved, no matter how plausible it looks.
Borel’s theorem is more than a century old. The simplest questions it raises about specific numbers are still open: is normal? Is normal? Is normal? All three are widely believed to be normal, but nobody has found a way to prove it.
Footnotes
-
D. D. Wall, “Normal Numbers”, PhD thesis, University of California, Berkeley (1949) – the source of the normality/equidistribution equivalence. ↩
-
John E. Maxfield, “Normal k-tuples,” Pacific Journal of Mathematics 3, 189-196 (1953) – the dependent-base case. ↩
-
The independent-base existence result is due independently to J. W. S. Cassels and to Wolfgang M. Schmidt, “Über die Normalität von Zahlen zu verschiedenen Basen”, Acta Arithmetica 7, 299-309 (1962). ↩
-
Émile Borel, “Les probabilités dénombrables et leurs applications arithmétiques”, Rendiconti del Circolo Matematico di Palermo 27, 247-271 (1909) – the original theorem. ↩
-
Tanguy Rivoal, “On the Bits Counting Function of Real Numbers”, Journal of the Australian Mathematical Society 85(1), 95-111 (2008) – the abstract states explicitly that simple normality in base 2 for , , and is conjectural. ↩
-
D. G. Champernowne, “The Construction of Decimals Normal in the Scale of Ten”, Journal of the London Mathematical Society 8, 254-260 (1933) – the original construction and proof. ↩
-
Wacław Sierpiński, “Démonstration élémentaire du théorème de M. Borel sur les nombres absolument normaux et détermination effective d’un tel nombre”, Bulletin de la Société Mathématique de France 45, 125-132 (1917) – the first explicit construction, though the construction did not provide an algorithm for computing its digits. A computable reformulation was later given by Verónica Becher & Santiago Figueira, “An Example of a Computable Absolutely Normal Number”, Theoretical Computer Science 270(1-2), 947-958 (2002). ↩
-
Yann Bugeaud, “Nombres de Liouville et nombres normaux”, Comptes Rendus de l’Académie des Sciences 335, 117-120 (2002) – proves both that normal Liouville numbers exist and that Liouville numbers normal in no base exist. ↩
-
Peter Trueb, “Digit Statistics of the First 22.4 Trillion Decimal Digits of Pi” (2016) – the digit-frequency computation cited above. ↩
-
Charles Hermite, “Sur la fonction exponentielle,” Comptes Rendus de l’Académie des Sciences 77 (1873), 18-24, 74-79, 226-233, 285-293. ↩
-
Ferdinand von Lindemann, “Über die Zahl π,” Mathematische Annalen 20 (1882), 213-225. ↩
-
F. M. S. Lima, “Some Transcendence Results from a Harmless Irrationality Theorem” (2014), lists among the numbers for which not even an irrationality proof is known. ↩