# The Curious Case of Learning Without Showing It
## A Puzzle in Plain Sight
In the world of machine learning, there’s a rule of thumb that everyone follows: if a model stops improving, it’s done. Training dashboards are watched like gauges on a dashboard, and the moment the numbers flatten, the experiment is considered finished. But what if that flat line is actually a period of invisible, underground construction?
This question came to a head in a deceptively simple experiment. A group of scientists set out to teach a very small artificial neural network to do modular addition — the kind of arithmetic you use on a clock, where the numbers wrap around after reaching a certain limit. Ask it what seven plus nine equals on a twelve-hour clock, and the answer is four. The task is trivial by human standards, yet it touches on a deeper question about how learning actually works.
The network absorbed the training examples at lightning speed. Within a few hundred rounds of practice, it could repeat every sample it had been shown with perfect accuracy. To any observer checking its progress at that stage, the experiment was a success. The model had learned.
Then the researchers pulled out entirely new problems — combinations the network had never encountered before. Its performance collapsed. It answered barely better than chance, as if it had been reciting a list of specific answers rather than grasping the underlying rule. This would have been the kind of result most papers would file away as textbook overfitting: the model memorized rather than understood.
But the researchers did something unusual. Instead of stopping there, they kept pushing the training far beyond the point where it appeared to matter. They let it run for thousands, then tens of thousands of additional steps — well past what any sensible early-stopping protocol would permit.
And then, at some point that no one predicted, everything changed. The network’s ability to solve unfamiliar problems surged upward, eventually reaching near-perfect accuracy on both the examples it had seen and the ones it had never encountered. The same training process, the same data, the same setup — but the model had quietly transformed into something fundamentally different inside.
The team gave this phenomenon a name: grokking.
## The Invisible Construction Phase
What makes grokking so unsettling is not the effect itself, but the fact that it happens during a period that looks like nothing at all. If you plot the model’s accuracy on new problems over time, you see a long, flat plateau followed by a sharp, almost discontinuous leap. Nothing on the surface suggests that a revolution is brewing underneath.
Imagine watching someone study for a certification exam. For weeks, they flip through flashcards and score perfectly on practice quizzes drawn from a bank of known questions. Their performance looks solid. You’d never suspect that they haven’t actually internalized the material. Then one day, you give them a question that uses the same concepts but in a new context, and they answer it effortlessly. The understanding wasn’t absent — it was under construction, invisible from the outside.
That’s the essence of grokking, but compressed into an extraordinarily short timeframe and made measurable with mathematical precision.
## Inside the Black Box: A Surprise Waiting to Be Found
What was the network actually doing during those thousands of apparently idle steps? A deeper investigation conducted by a separate team of researchers revealed a striking answer. Rather than just brute-forcing its way through patterns, the model had essentially reinvented a branch of mathematics.
It learned to place numbers at specific points around a circle — much like the positions of hours on a clock face. Once numbers had these spatial representations, addition became a geometric operation: rotate from the position of the first number by the distance corresponding to the second number, and read off the result. The network had independently discovered the connection between circular geometry and cyclical arithmetic, a relationship formalized through trigonometry centuries ago by human mathematicians.
Nobody programmed this. No human designer guided the network toward this solution. It emerged on its own, driven purely by the optimization process, as if the algorithm had found a more elegant and general way to solve the problem than simply memorizing answers.
For a long time, two competing strategies coexisted within the same network. One was a fast but shallow memorization shortcut that worked only on familiar examples. The other was the slow, painstaking construction of a circular representation that promised true generalization. Only when the deeper method was fully assembled did it overpower the shortcut, and only then did the network’s performance on new problems suddenly take off.
## Why This Matters Beyond Clock Math
The experiment was conducted on a toy problem, with a toy model. But the implications ripple outward into larger, more consequential territory. Modern AI systems are trained for days, weeks, or months on enormous datasets, using automated procedures that routinely shut down training when progress appears to stall. If grokking can happen in a simple clock-arithmetic network in a matter of hours, what similar transformations might be occurring — unseen — in systems we rely on for language, reasoning, and decision-making?
This is not a call to abandon early stopping or throw caution to the wind. It’s an invitation to look more carefully at what “finished” really means. A plateau in performance is not necessarily evidence of a finished process. Sometimes it’s evidence of a process that hasn’t yet completed itself.
The broader takeaway is humbling. The surface-level metrics we use to judge learning — test scores, accuracy, loss curves — may only capture the tip of the iceberg. Beneath them, models can be developing representations, structures, and strategies that have no visible correlate until the moment they suddenly do. Understanding, it turns out, doesn’t always announce itself when it arrives.
## The Word Itself
The term “grokking” comes from science fiction writer Robert Heinlein, who used it in his 1961 novel *Stranger in a Strange Land*. In Heinlein’s vision, grokking meant understanding something so completely that it became part of you — not just knowledge you could recall, but knowledge you had absorbed so deeply that it was intuitive, instinctive, and inseparable from your way of thinking.
It is a fitting word for what the neural network experienced. At some hidden inflection point, the model stopped pattern-matching and started genuinely “knowing” the rule. The transition was sudden, decisive, and impossible to predict from the outside.
## FAQ
**Q: Is grokking limited to tiny networks and simple math problems?**
A: So far, the phenomenon has been most clearly documented in small models performing rule-based tasks like modular addition and other algorithmic puzzles. Whether grokking occurs in large, modern language models during the training of complex real-world tasks remains an active area of research. The question is still open.
**Q: Can we detect grokking before it happens to prevent wasting compute?**
A: Not reliably — which is part of what makes the phenomenon so intriguing. Researchers have tried various internal measurements and proxy signals to predict when the “click” will happen, but no method has proven fully effective yet. The internal structural changes that precede the performance jump are difficult to monitor in real time.
**Q: Does grokking mean the model is “conscious” or has “understanding” in the human sense?**
A: No. When researchers say the model develops a generalizable representation, they mean it has built an internal structure that correctly captures the underlying rule. This is a functional form of generalization, not a subjective experience. It is remarkable that such structures can emerge from optimization alone, but it is not the same as human understanding.
**Q: Is early stopping bad practice?**
A: Not at all. Early stopping remains a powerful and widely used technique to prevent models from overfitting and to save computational resources. The lesson from grokking is not to discard early stopping, but to recognize that a flat performance curve does not always mean “nothing more to learn.” Context and the nature of the task matter.
**Q: Has grokking been reproduced in other settings?**
A: Yes, the phenomenon has been observed in several different algorithmic and rule-based learning tasks beyond the original modular addition experiment. Each case reinforces the idea that memorization and generalization can coexist temporarily, with the transition between them being harder to predict than it first appears.
## Conclusion
Grokking challenges one of the most fundamental assumptions in machine learning training: that if a model isn’t improving, it’s done. The experiment with a simple clock-arithmetic network showed that genuine, generalizable understanding can remain hidden in plain sight — invisible on the performance graph, indistinguishable from mere memorization — until it suddenly isn’t.
This forces us to reconsider what we think we know about the learning process, whether in machines or, by analogy, in ourselves. The moments when we feel stuck, when progress has stalled and nothing seems to be happening, may be the very moments when the most important structural work is taking place beneath the surface.
The deeper lesson is about patience with complexity. Understanding doesn’t always follow a smooth, predictable path. Sometimes it arrives late, suddenly, and only after a long quiet period that, from the outside, looked like failure.
Thank you for reading



