To main content To menu

Grammar All The Way Down Then I found four pebbles

Ultrafinitism wants infinity banished from math - but every fix for infinity is just another grammar, and quantizing an 80B model shows why.
A lone figure seen from behind, crouched on a vast cracked plain, arranging four dark pebbles in a row on a stone slab while hundreds more scatter toward a hazy horizon under an enormous sky
Four stones counted, and the rest of them out there past where counting goes Image: AI-generated with google/nano-banana-pro by lui.vn

The monkey learns to count to three. Genuine achievement, that – three took a while, three cost him something. Then he finds a fourth pebble on the ground, turns it over in his hand, and concludes he has discovered infinity.

That monkey is me at 1 am, reading about mathematicians who want infinity abolished.

The heretic with the discrete necklace

Doron Zeilberger is a combinatorialist at Rutgers who thinks the universe ticks. Not flows – ticks. Where the rest of us see a smooth continuum, he sees a flip-book being turned fast enough to fool us, and he has spent decades saying so loudly enough that colleagues cross the room at conferences. His line, delivered to a BBC camera in 2010 and quoted back at himself ever since (he is fond of quoting himself, and honestly, fair): Infinity may or may not exist; God may or may not exist – but neither of them, he says, belongs in mathematics.

He is an ultrafinitist. The position goes further than plain finitism, which merely refuses completed infinities. Ultrafinitism says a number exists only if you could actually make it – write it, compute it, verify it with real resources on real hardware in real time. Which puts a ceiling on the number line, somewhere.

The ceiling is where the universe comes in, and it is why I got pulled into this in the first place. There are ≈10^80 atoms in the observable universe. The covariant entropy bound caps the information content of a causal diamond in a universe with our cosmological constant at ≈10^122 bits. Seth Lloyd worked out in a 2002 paper that the universe has performed at most ≈10^120 elementary logic operations on ≈10^90 bits since the Big Bang (≈10^120 bits if you count gravitational degrees of freedom, which is a fun sentence to type). Past that budget, nothing physical instantiates anything. So take Skewes’ number, e^e^e^79, which nobody has ever written in decimal and nobody ever will. Is it odd or even? Is it prime? Could you find one anywhere in nature, ever? Zeilberger’s answer is that maybe it simply isn’t a number.

I read that and thought: obviously wrong. And also, uh, wait.

Almost every number has no name

The thing that actually broke my brain has nothing to do with cosmology. It is a counting argument you can do on a napkin, in about four lines, with a coffee going warm next to it.

Any description of a number is a finite string of symbols. An algorithm, a formula, a definition, a sentence in German, a poem, a Python function – all of them finite, all of them built from a finite alphabet. And there are only countably many finite strings. You can put them in a list: all the one-character strings, then all the two-character ones, and so on forever. The list is infinite but it is countable, and countable means it has the same size as the whole numbers.

The real numbers are not countable. Georg Cantor proved that in 1891 with an argument short enough to fit on a napkin next to the first one.

Put those together and you get something worse than uncomfortable. There are only countably many nameable numbers and uncountably many numbers, so almost every real number cannot be described. Not “hasn’t been described yet”. Cannot – by any finite string, in any language, ever, by anyone, including languages nobody has invented. The nameable reals are a measure-zero cameo inside the reals. Pi and the square root of two are not typical numbers. They are freaks, and they are freaks precisely because they have names.

Which quietly dissolves the objection I started with: you cannot compute pi exactly – true, and irrelevant. Pi’s information content is not its decimal expansion, it is the recipe, and the recipe is tiny. There is a formula that hands you the nth hexadecimal digit of pi without computing any of the previous ones (the Bailey-Borwein-Plouffe formula, and yes, it feels like cheating). Nobody was ever storing the digits, the same way nobody stores a river. Saying pi doesn’t exist because the expansion never finishes is like saying a song doesn’t exist because you haven’t finished singing it.

The scandal is one floor down, and it is not about the numbers we can name – it is about the ones we can’t, which is nearly all of them.

Skolem’s joke

In 1922, Thoralf Skolem noticed something that should be more famous than it is: if standard set theory has a model at all, it has a countable one. A universe of sets containing only countably many objects, which nonetheless proves – correctly, internally, with full rigor – that the real numbers are uncountable. Both things are true at once. The bijection that would collapse the reals into a countable list does exist; it just lives outside the model, where the model can’t reach it.

So “uncountable” turns out not to be a statement about how many things there are. It is a statement about what the language can see from where it is standing.

That is Skolem’s paradox, it has been sitting in plain view for a century, and it is the most precise version I know of the thing that has been nagging at me. Nine years later Kurt Gödel showed the language can’t prove everything true about its own subject. Two years after that, Tarski showed it can’t even define its own truth predicate. Three theorems, three decades, one finding: the grammar bounds what you can say, and the boundary is not where you’d like it.

The grammar is the limit. It is such an obvious sentence that it takes a while to notice it has teeth.

The exit is also a room

Which brings me back to Zeilberger, who did not survive my own argument either.

Ultrafinitism presents itself as the way out. Stop assuming infinity, work only with what can be built, and the paradoxes evaporate along with the mysticism. Replace the continuous line with a discrete necklace of numbers separated by tiny-but-not-infinitesimal gaps, rewrite calculus as difference equations, let the computer do the ugly parts. Zeilberger lists his own machine as a co-author on his papers; the man is committed.

But look at what he has actually done. He has swapped one set of conventions about what counts as a legitimate object for another set of conventions about what counts as a legitimate object. That’s not an escape from grammar – that’s a different grammar, with its own load-bearing assumption sitting right where the old one sat.

And this one is vaguer. Where exactly is the cutoff? Nobody can say, and most ultrafinitists hold that the line between feasible and infeasible is inherently fuzzy – which is a philosophical position, not a dodge, but it does mean the movement’s central object is undefined on purpose. When Harvey Friedman pinned Alexander Esenin-Volpin down and walked him up the powers of two – does 2^1 exist? 2^2? 2^3? – Esenin-Volpin said yes every time, but took proportionally longer to say it each time. Beautiful bit. Perish the thought that it’s a theory.

Somebody did try to make it a theory. Edward Nelson woke up one morning in 1976 at Princeton convinced that the infinite world of numbers he’d believed in was an arrogance, and set out to rebuild arithmetic without it. What he got back was remarkably weak. Axioms that banished infinity couldn’t prove that a + b equals b + a. Exponentiation stopped being guaranteed – you might construct 100, or 1,000, and then fail to construct 100^1000. Induction was gone entirely. In September 2011, after twenty-five years of work, he announced on a mailing list that he had proved standard arithmetic inconsistent. Terence Tao and Daniel Tausk independently found the error within days. Nelson withdrew the claim on October 1, thanked them, and went back to work the next morning – which is, whatever else you think of him, exactly how it’s supposed to go.

So: the grammar that promises to fix the grammar problem cannot count reliably. Delete infinity and you lose commutative addition. That is the cost, itemized, by the person most motivated to find it cheap.

Twenty-eight bits nobody was using

I would have left it there as an interesting evening if I hadn’t spent last month doing the ultrafinitist experiment by accident, on rented hardware, for entirely unrelated reasons.

I wanted to understand model quantization, which had been a black box to me – the kind you use daily and can’t open. Well. So I rented B200 time at Spheron and quantized Qwen3-Coder-Next, an 80,000,000,000 [80 billion] parameter model, symmetric and asymmetric, down through the bit-widths. Then I ran the results locally on Bazzite with an RTX 5080 and 64 GB of RAM, through ollama, and measured what fell out.

Quantization is precision demolition. You take weights stored as 16-bit floats and re-encode them in 8 bits, then 4, then 2, and each step throws away information you cannot get back. The speed payoff is immediate: bfloat16 gave me ≈30 tokens/s, 8-bit symmetric ≈45, and both 4-bit variants cleared ≈60, roughly double the baseline.

The quality question is the interesting one: I scored each version on GPQA Diamond, the hard 198-question core of a benchmark built to be Google-proof: graduate-level biology, chemistry and physics, four options each, so a coin-flipping idiot lands at 25 %. On the full 448-question set, PhD-level domain experts scored 65 % and skilled non-experts with unrestricted web access and half an hour per question managed 34 %. On the Diamond subset specifically, PhD experts recruited by OpenAI later scored 69.7 %.

bfloat16 scored 64.9 %. A rounding error under the human expert line, which I did not expect from a coder model being quizzed on chemistry. (Numbers like this depend on the eval setup – the same weights can move several points between zero-shot and chain-of-thought settings. Mine are one setup, one run, one guy).

8-bit symmetric scored 71.4 %. Higher than the full-precision original – though on 198 questions that gap is around 13 answers and well inside the noise a sample that size can resolve, so the honest reading is “at or slightly above baseline”, not “compression improves the model”.

4-bit symmetric scored 64.4 %. Read that again. I threw away 28 of 32 bits, doubled the throughput, and the model got exactly one more PhD-level science question wrong than it had at full precision. One.

4-bit asymmetric came in lower at 60.8 %, which is backwards from what the theory says should happen at low bit-widths, and I have no explanation for it. My method was amateur (I’m a curious bored guy, not a scientist, and the difference showed).

Then 2-bit asymmetric scored 1 %. Below random guessing. A model that scores under chance on multiple choice isn’t a degraded model, it’s a model that has stopped emitting parseable answers at all, which matched what I saw when I poked it by hand: it refused, or produced mush. I also ran five questions of my own past every version – the old name of Ho Chi Minh City, the country whose capital is Pyongyang, 6 + 7, the common name for H2O, and the difference between a Krapfen and a Berliner. Everything down to 4-bit answered all five. The last one has no correct answer and I asked it as a German, in bad faith, on purpose.

Either 2-bit is too much information loss for this model, or my method was too naive to measure it properly. The truth is probably in between.

What the machine says when it runs out of room

Here is what quantization taught me that no amount of reading about Skewes’ number did: the cliff is real. It is nowhere near where the notation implies. Somewhere between 4 bits and 2 bits, a language stops being able to say the thing – and above that line, the extra precision was largely decoration, doing far less work than the format’s existence suggested it was doing.

And the failure mode, when it comes, is not gradual fuzz. It is Infinity, arriving like a fuse blowing. Drop to float16 and your maximum representable value is ≈65,504; overflow past it and the number doesn’t get big, it gets replaced by a symbol meaning “off the end of the alphabet”. Then Infinity minus Infinity gives you NaN, and NaN propagates through every operation it touches, so one bad activation and the whole tensor is fucked.

Which tells you exactly what infinity is inside a computer: not a quantity. A road sign at the edge of the map. IEEE 754 needed a token for “this exceeded the grammar” and picked the most loaded word in mathematics to do it. Bold choice!

The fix is instructive too. bfloat16 uses the same 16 bits as float16 and simply reallocates them – fewer to the fraction, more to the exponent – and the infinities go away. Same information budget, different distribution, different set of expressible worlds. The overflow was never a fact about the numbers, but a fact about the container.

JavaScript, meanwhile, ships Number.POSITIVE_INFINITY as a literal, a named constant standing in for a concept the machine cannot represent. Infinity + 1 returns Infinity. Cantor’s absorption arithmetic, live in production, in a props type, in code I have shipped without a second thought (and so have you, don’t pretend otherwise).

The languages all agree, and that’s the problem

Nicolas Gisin, a quantum physicist in Geneva, has spent years arguing that real numbers are the hidden variables of classical physics – that when we write an initial condition as a real number, we are quietly assuming infinite information was determined at the start of time, and then acting surprised that the future comes out looking fixed. Rewrite physics in Brouwer‘s intuitionistic mathematics, where numbers are unfinished processes that acquire digits as time passes, and indeterminism appears without a single equation changing. Real numbers are not really real, as he keeps putting it.

The part of his argument that unsettles me most is the concession. He is explicit that both mathematical languages agree on every computation a physicist will ever run. Identical predictions, down the line. What differs is only the picture of the world the language pushes into your head while you use it.

So determinism might never have been a discovery. It might be a grammatical feature we inherited and mistook for a finding.

And ultrafinitism – the movement that spotted this exact problem, that saw infinity smuggled in through notation and objected – answers it by smuggling in a cutoff nobody can locate. Three languages. Zeilberger’s discrete necklace, Brouwer’s unfinishable continuum, standard set theory’s completed infinities. Same predictions, three incompatible worlds, no experiment between them.

Nobody gets out. Least of all the people selling the exit.

Four pebbles

I keep going back to that monkey, because the story is worse than I first told it.

He counts to three. He finds a fourth pebble and names it infinity. Then he builds a language sturdy enough to name the fourth pebble properly, and the language works so well he builds all of physics on it. Then he proves – inside that language, using that language – that it cannot name almost any of the pebbles. Then a heretic offers him a smaller language with no pebbles past three in it at all, and the heretic can’t say where three ends, and in the new language you can’t prove that two plus two equals two plus two.

Meanwhile the pebbles just sit there in the dirt, warm, entirely unbothered.

Four of them. Still four.