
Around 300 BCE in Alexandria, Elements, humanity’s second best seller, opens with a definition, ‘a point is that which has no part’, followed by ‘a line is a breadthless length’. Euclid compiled three centuries of geometric discoveries of early Greeks and ordered them through a single series of impressive deductive links. For the next two thousand years, the book became the shoulders upon which many great thinkers stood to see farther and expand the repository of our knowledge of the world. Archimedes, Isaac Newton and Abraham Lincoln were among many who studied Euclid.
Euclid had the considerable advantage of being right by design. Each theorem is compelled by what precedes it, ultimately built on axioms that appear plain and evident on their own. His ability to derive an entire mathematical world through logic alone captivated humanity for over two thousand years. Today, we still teach Euclid to schoolchildren as the way space is. For example, the three angles of a triangle sum to 180 degrees; parallel lines never meet; a2+b2=c2 in a right triangle, etc. Proofs built upon axioms also conditioned younger me to take what I was taught in classes as both sophisticated and certain throughout my teens. They felt as evident as ‘The capital of France is Paris’.
Of course, Einstein had something to say about this. In 1921, five years after publishing his theory of general relativity, he gave a lecture in Berlin. There, he expressed one of the sharper scepticisms towards our familiar belief system derived from Euclidean mathematics. He asked, ‘how can it be that mathematics, being a product of human thought which is independent of experience, is so admirably appropriate to the objects of reality?’ He acknowledged the certainty of the Euclidean laws, but drew the line that cuts the knot: ‘as far as the propositions of mathematics are certain, they do not refer to reality’. Axioms are ‘the creations of the human mind’: the point, the straight line. The irony implied is that geometry, a discipline whose name means earth-measuring, had spent two thousand years without consulting the real earth. For example, three angles of a triangle will exceed 180 degrees if drawn on a sphere; parallel lines can meet in curved space; and the Pythagorean theorem fails on a curved surface. After all, Euclidean geometry requires a strictly flat space, and flat space is only an extremely localized patch of our reality. Axioms bound by a localized reality cannot represent the universal truth. Any attempt at the universal truth must surrender certainty.
A similar challenge arises in discussing intelligence. Like geometry, intelligence is not confined to a single, universal framework. Intelligence is an amorphous quality, and we can only infer its presence from behaviour. In fact, behaviour is all we ever observe of the doer underneath. These inferences, like Euclid’s axioms, are inevitably local. The evolutionary record also points to this. Roughly 600 million years ago, a small cluster of neurons evolved at the end of a worm-like animal. It did one job: steer. It turned a body towards food and away from harm. It did not model its surroundings or plan any routes. Many scientists point to this milestone as the starting point of neural evolution that we associate with intelligence. Since then, nature has witnessed the evolutionary lineage of intelligent systems diverge and converge along the way, branching out from first worm-like animals into vast life forms such as whales, honeybees, birds, octopuses, and of course, humans. The intelligence manifested within each species, including humans, has evolved to have its niche and features, but is by no means exhaustive nor universal.
Artificial Intelligence is no exception. Although its features and depth have expanded radically fast in recent decades, it still represents a localized version of intelligence. AI illustrates a unique way of processing information and forming outputs that is foreign to biological forms that are known to us. The difference is, however, that the inputs of its learning are built upon human texts and reasoning, and the outputs are designed to emulate human behaviour. While we acknowledge intelligence in other life forms, we rarely put much weight on the presence of intelligence in those during our day-to-day lives. However, the intelligence manifested through AI captivates us with both awe and fear, not because the intelligence is merely present, but because we see a shadow of our own version of intelligence without being us.
Awe and fear are both rooted in the encounter of the unknown. Both emotions arise when we face something beyond our understanding or ability to control. When IBM’s Deep Blue beat Garry Kasparov in 1997, the world was shocked. It was the first time that AI showed a superhuman performance in a game that long signalled human intellectual superiority. However, shortly after the AI’s victory, The New York Times published an article demystifying the brute-force calculator behind Deep Blue’s façade. By explaining its architecture and calling out its limitations, many of us were reassured that the super-calculator was not a thinking machine. In the same article, the New York Times also moved the goalpost of AI’s cognitive capability to the game of Go. When DeepMind’s AlphaGo beat Lee Se-Dol nineteen years later, we saw a similar psychological cycle repeat. We had spent nearly two decades believing that Go required human intuition and judgment that the brute-force calculator could not reach. While this belief did not survive AlphaGo’s victory, we also soon internalized that AlphaGo was a specialized system built upon reinforcement learning, but largely powerless outside of Go at the time. Each time, within months, we formed an abstract framework to locate the boundary of a model’s capability. Then, in 2017, Google researchers published the Transformer architecture that made it practical to train models on massive amounts of data, which enabled the birth of ChatGPT, Claude and Gemini. This time, the demystification never arrived.
Until LLMs, we were given a framework to draw a boundary of each intelligence system. Deep Blue’s demystification came about because the account of its design covered the full scope of its behaviour; 200 million chess positions per second through brute-force computing. Despite the opacity of its neural network, we became comfortable with AlphaGo because Go had a deterministic rulebook, and AlphaGo’s superhuman ability remained within its 19×19 grid. However, LLMs inverted this pattern. The most common explanation of their design is simple: ‘it predicts the next token’. Unfortunately, this provides near-zero transparency about what it can do. At the same time, its capability goes beyond any conventional grids, spanning coding, poetry, and the expression of empathy, with a voice of its own.
Nevertheless, the LLM still represents one more locality of intelligence. Although the opacity of LLMs makes it difficult to draw a boundary around their capability the way we could with Deep Blue or AlphaGo, a boundary we cannot see is still a boundary. Therefore, our awe and fear require recalibration to acknowledge the locality of its competence, just like every other intelligence before it. What we need is a better map that can help measure AI’s capability more precisely and more quickly.
Today, OpenAI and Anthropic spend heavily on AI safety. However, much of their public-facing work has focused on identifying the unintended behaviours of their own systems. Their published findings disproportionately expand the catalogue of behaviours AI is capable of but was not intended to exhibit, while doing much less to establish where AI reliably fails. But the reverse is equally important. Establishing where AI reliably fails and tying those limits back to social and economic implications is what demystifies AI for the public. Furthermore, the access to capability knowledge still belongs to insiders. As a result, whatever demystification comes along the way becomes inward-facing. As the understanding accrues inside, the demystification cycle for the rest of us stays delayed. Without the map, those outside the frontier labs are left with less agency to write their own future with AI. These stakes compound. As the next layer of machine intelligence arrives, the cost of measuring it, as well as the share of risk we cannot understand, will be greater than the time before.
The good news is that there are already a few great surveyors for the map. METR measures the autonomous capabilities of AI; Epoch AI tracks the driving forces (compute, costs, data) behind AI. Transluce publishes public-facing frameworks to measure AI’s behavioural propensities. The bad news is that (1) these organizations are too small and fragmented, (2) their findings are hard to translate for the general public because the work is technically dense, and (3) most importantly, there is no cartography that can confidently point to what the next evolution of AI will entail. As a result, when OpenAI says they are 80% of the way to AGI, we have little idea what capabilities they are referring to, or whether it’s just another marketing campaign. And when the Anthropic CEO predicts a ‘country of geniuses in a datacenter’ by 2027, we are still left wondering where we are, how we will get there, and what we should observe along the way: a destination with no position, route or milestones.
The truth of Euclidean geometry turned out to be local. But the Elements laid the logical scaffolding that made the beyond-local accessible and comprehensible. Einstein called it his ‘holy little geometry book’. Euclid’s method became the foundation of what came next. Our cartography of AI may meet the same fate. Any specific measurement of today’s AI will be superseded when a new model arrives. But our method will compound. Like the Elements, it will become the foundation for understanding what comes next. As new waves of intelligence arrive, we will be able to meet each one knowing more than we did at the last.
Leave a comment