# Can Artificial Intelligence Recreate the Greatest Scientific Discoveries of History?
The year 1915 marked a turning point in humanity’s understanding of the universe. A physicist working in isolation published a framework that redefined gravity not as a force but as a warping of the very fabric of space and time by mass. That framework — general relativity — went on to shape modern cosmology, predict black holes, explain gravitational waves, and even guide the satellites in our GPS systems. It stands as one of the most profound intellectual achievements in human history.
Now, researchers in artificial intelligence are asking a bold question: could machines ever replicate such a feat? And what would it take for an AI to produce a breakthrough of that magnitude?
## The Einstein Test
The idea of testing AI against the greatest scientific milestones gained significant attention earlier this year at a major summit on AI held in New Delhi. One of the leading figures in AI research proposed a fascinating experiment: train a large language model exclusively on knowledge available before 1911, and then see whether the system could independently arrive at the principles of general relativity. Such a result, he argued, would constitute a meaningful benchmark for what researchers call artificial general intelligence — the elusive goal of building machines capable of human-level reasoning and discovery.
The appeal of this idea is intuitive. If an AI trained only on pre-1911 data could derive a theory that fundamentally reshaped our understanding of physics, that would suggest a capacity for genuine insight far beyond pattern recognition.
But the idea has since taken on a life of its own. In late 2024, a researcher at a nonprofit focused on AI alignment gave a talk exploring what he called “vintage” language models — systems trained only on historical text up to a specific date — and what such models might be able to rediscover. Since then, several independent teams have begun building working versions of these historically constrained models, hoping to push the boundaries of what AI can achieve in scientific reasoning.
## The Nature of Breakthrough Thinking
What makes Einstein’s achievement so difficult to replicate — and so instructive — is the type of reasoning it required. Most AI systems today excel at inductive reasoning: they identify patterns across enormous datasets and produce the most statistically probable outputs. This makes them exceptional at tasks like translation, summarization, and even solving well-defined mathematical problems.
But scientific revolutions rarely emerge from the accumulation of data alone. Philosophers of science have long distinguished between induction and a deeper form of reasoning called abduction — the art of inventing a cause for a singular or puzzling phenomenon. A breakthrough like general relativity is not a matter of finding the most likely explanation given existing data. It is a creative leap, a paradigm shift, that emerges when the existing framework is shown to be inadequate.
Historian of science Thomas Kuhn described these moments as paradigm shifts — times when the very assumptions underlying a scientific discipline are overturned. These shifts frequently happen not when data is abundant, but when it is scarce, contradictory, or anomalous. A developing model of the solar system, combined with meticulous observations of planetary motion, eventually led Newton to formulate his law of universal gravitation. The key insight was not statistical — it was conceptual.
The central challenge for AI, then, is whether current models can develop something akin to a “world model” — an internal representation of underlying principles — from sparse and incomplete information. Can they reason their way from a handful of clues to a general framework that explains the whole?
## Building the Vintage Machines
After the Einstein test was proposed, an independent AI researcher based in San Francisco set out to try something similar. He built a language model he called Machina Mirabilis, trained exclusively on texts published before 1900, and tested whether it could independently produce the theories behind quantum mechanics, special relativity, and general relativity.
He provided the model with helpful nudges along the way. In one instance, he fed it observations about the photoelectric effect — the phenomenon in which light dislodges electrons from metal surfaces — a clue that would eventually lead to the quantization of light. In another, he described Einstein’s famous elevator thought experiment, in which a person in a sealed elevator cannot distinguish between the effects of gravity and acceleration.
The results were a mixed bag. In some cases, the model showed tantalizing flashes of something resembling intuition. When presented with the photoelectric effect, it described light as breaking “into a multitude of distinct impulses” — a hint that gestured toward the concept of light quanta. But in most cases, the model failed. It lacked genuine understanding of the physics it described, sometimes producing plausible-sounding language without any coherent internal model of the world to reason from. One observer noted that the system was “parroting words that seem plausible” without any deep grounding.
Training a model on such historically constrained data also proved enormously difficult. The texts from that era contain broken English, archaic terminology, OCR errors from digitization, and inconsistent formatting — making it hard to create a clean, reliable training set.
Another team, led by an independent computer scientist, attempted a similar project with a cutoff date of 1930. The choice was practical: works published in 1930 entered the public domain in the United States at the start of 2026, making the legal landscape simpler. This team imagined using such a model to investigate questions about discoveries from that era — Turing machines, Gödel’s incompleteness theorem, or the theoretical prediction of neutrinos. But they encountered the same data problems. Without careful dating of training materials, the model would inadvertently “leak” information from later decades. In one revealing test, it answered questions about the administration of a mid-twentieth-century president — a person who was not even in office during the supposed cutoff period.
Despite these failures, one team member expressed optimism. He envisioned a model trained only on data up to a recent date and asked to generate predictions — essentially creating a “forecasting engine” for scientific discovery. He believed that with sufficient scale, such a model would eventually make a non-trivial discovery.
A separate group at a Swiss university has developed an entire family of historical models under a project called Ranke, with training cutoffs at 1913, 1929, 1933, 1939, and 1946. Because they lack the computing power and data volume of the largest modern AI systems, their ambitions are more modest. They are not searching for genius, they say, but rather for “sparks of genius” — small seeds of ideas that might echo discoveries that came later.
## The Problem of False Theories
Interestingly, there is no obvious technical barrier preventing a language model from producing a correct scientific theory. In principle, a probabilistic system might generate the general theory of relativity as one output among thousands of incorrect alternatives. The real difficulty lies in knowing which output is actually worth testing.
One team found this out firsthand when they gave a model a foundation in orbital mechanics and synthetic data mimicking various planetary systems. The model never arrived at the true law of gravitation — the one relating force to mass and distance. Instead, it inferred a different, incorrect law for each planetary system, each wrong in its own unique way. The model produced internally consistent but completely false explanations.
This highlights a fundamental difference between how humans and AI systems devise theories. When scientists develop competing hypotheses, they rarely do so by assigning probabilities to a vast range of possibilities and then searching for the correct one. They build theories from conceptual leaps, guided by intuition, elegance, and the sense that a particular framework feels right even before it can be fully tested. Current AI models, by contrast, operate through statistical prediction — and this approach is ill-suited to the kind of imaginative reasoning that drives the greatest scientific advances.
There is a silver lining, however. In recent months, AI systems have made striking progress in pure mathematics. One notable achievement involved an AI system that disproved an eighty-year-old mathematical conjecture posed by a famous Hungarian mathematician. Experts described the solution as a “genuine conceptual advance,” noting that it required small, novel abstractions along the way — not just brute-force searching. At the same time, the solution drew on ideas already present in the existing mathematical literature, so it was not quite an Einstein-like flash from nowhere. But it demonstrated that AI can sometimes contribute something genuinely new and insightful within a well-defined domain.
## Where Do We Go From Here?
Researchers agree that a relativity-like breakthrough for AI is not inherently impossible. Several experts have argued that achieving such a feat would require rethinking the foundational principles on which today’s models are built. The current architecture — optimized for pattern matching and statistical prediction — may simply not be the right tool for abductive, creative reasoning.
The vintage model experiments have been illuminating not because they have succeeded, but because they have revealed just how far current AI is from genuine discovery. They show that while large language models are increasingly powerful tools for exploration and synthesis, the kind of reasoning that produced general relativity — the ability to leap beyond the available data, to imagine a new framework from sparse clues, to see the world differently — remains stubbornly beyond their reach.
But perhaps the most valuable lesson is a humbling one. These experiments remind us that scientific genius is not just about having access to information. It is about the capacity to organize that information in ways no one has imagined before — a capacity that, for now, belongs to human minds alone.
—
## Frequently Asked Questions
**What are “vintage” or “historical” language models?**
These are AI language models that are trained exclusively on text from before a specific cutoff date — for example, all knowledge available before 1911 or before 1930. The goal is to see whether the model can rediscover scientific ideas or make predictions using only the information that was available during that historical period.
**Why is general relativity used as a benchmark for AI?**
General relativity is one of the most consequential scientific discoveries in history. It required a conceptual leap that went far beyond simply finding patterns in existing data. Proposing that an AI could recreate this discovery — using only pre-1911 knowledge — serves as an ambitious test of whether current AI systems can achieve a level of reasoning comparable to human genius.
**What is the difference between inductive and abductive reasoning?**
Inductive reasoning involves drawing general conclusions from specific examples and data patterns — the kind of reasoning at which AI excels. Abductive reasoning is the process of forming the best possible explanation for a given observation, often involving creative leaps and conceptual innovation. Scientific breakthroughs like general relativity rely on abductive reasoning, not inductive pattern matching.
**Have any vintage AI models successfully recreated a major scientific theory?**
No. While some models have shown glimpses of promising reasoning — such as hinting at the concept of light quanta when given observations about the photoelectric effect — none have succeeded in independently producing a full, correct scientific theory like general relativity. The results so far highlight the limitations of current AI architectures rather than their strengths.
**Why is it so difficult to train a historically constrained model?**
Training data from earlier centuries is riddled with OCR errors from digitization, archaic language, inconsistent formatting, and inaccurate timestamps. This makes it very difficult to ensure the model is truly limited to pre-cutoff knowledge. In several experiments, the models “leaked” information from later periods, sometimes answering questions correctly about events that had not yet occurred in the assumed training timeframe.
**Can AI make genuine scientific discoveries?**
AI has already contributed to mathematical research in meaningful ways, including disproving long-standing conjectures through novel abstractions. However, these achievements draw on existing knowledge and established frameworks. A truly transformative discovery — one that creates an entirely new paradigm, as general relativity did — remains beyond the current capabilities of AI, though researchers are actively working to push these boundaries.
—
## Conclusion
The experiments building “vintage” AI models to recreate historical scientific breakthroughs have yielded more questions than answers — and that may be precisely their greatest value. They have exposed the fundamental limitations of current AI architectures in performing the kind of creative, abductive reasoning that drives the most important discoveries in science. While large language models continue to grow in capability and are becoming increasingly useful as tools for exploration and synthesis, the leap from pattern recognition to paradigm-shifting insight remains a uniquely human capacity — for now. As researchers refine their approaches and rethink the principles underlying AI design, these early attempts serve as both a benchmark and an inspiration: they remind us of what general relativity accomplished and of the extraordinary distance between statistical prediction and true understanding.
Thank you for reading



