# The One-Way Fact Problem: Why Language Models Struggle to Answer Questions From the Other Direction
Imagine you’re introduced to someone named Clara, and you learn that Clara is a celebrated violinist. Later, at a concert, a friend points to a woman on stage and asks, “Who is that?” If you recognize her, you can answer from either direction — you knew Clara was a violinist, and now you see the violinist and recognize Clara. This kind of flexible, two-way knowledge is so basic to human thinking that we barely notice we’re doing it. But it turns out that language models don’t always share this ability.
There’s a well-documented quirk in how these systems learn and store information: when a model is trained on a fact written one way, it often fails to recall that same fact when the question is phrased from the other direction. A model that learns “Clara is a celebrated violinist” through training may answer confidently when asked “Who is Clara?” but stumble or guess blindly when asked “Who is a celebrated violinist?” This one-way blind spot has become known as the Reversal Curse.
## The Surprising Asymmetry
At first glance, the asymmetry seems absurd. “Clara is a violinist” and “A violinist is Clara” carry the same meaning. If a model can derive the meaning from one phrasing, surely it can work backwards?
The key distinction is between *reasoning with a fact that’s sitting right in front of you* and *recalling a fact you learned during training*. When the fact appears in the prompt itself, even a large model can often work out the reverse. The trouble surfaces when the model has to pull the fact from memory, from the weights it adjusted during training. And in those stored weights, the direction matters enormously.
This has been demonstrated across multiple scales. Large language models fine-tuned on pairs of related entities tend to perform extremely well — sometimes nearly perfectly — when the question matches the training direction. But flip the question around, and accuracy can collapse to near chance. In one set of tests, a model could name a celebrity’s parent about four out of five times, but when given the parent’s name and asked for the celebrity, it only succeeded about one in three times.
## A Toy Model You Can Build in an Afternoon
To really understand what’s going on, it helps to strip the problem down to its bare minimum. The question becomes: how small and simple can a model be and still exhibit this one-way blind spot?
The answer might surprise you. It doesn’t take a billion-parameter transformer or months of compute. A model you could write from scratch in a few hours, using nothing more than basic array operations, is enough to show the effect clearly.
This toy model works like this: every word gets turned into a short list of numbers — think of it as the model’s private notes about that word. These notes start out random and get adjusted each time the model encounters a word in a particular role. To predict the next word, the model simply adds together the notes for the words it just read and produces a score for every word in its vocabulary. The word with the highest score is its guess. There’s no attention mechanism, no hidden layers, no transformer architecture — just two small matrices and a straightforward prediction rule.
The deliberate simplicity is the point. By removing everything a modern language model has, we can isolate exactly what single-directional training does on its own, without the noise of scale, pretraining, or architectural complexity.
## Setting Up the Experiment
The test is elegant in its design. Two hundred made-up facts are created, each pairing two unique names that appear nowhere else. For example: “Xyphos Brin is the Keeper of Echoes.” Because these names are entirely invented, the model can’t rely on any prior knowledge — everything it learns comes exclusively from the sentences it’s shown.
For each fact, a coin is flipped to decide which direction to teach it in. Half the facts are presented as “Name A is ___,” training the model to fill in the description that follows. The other half are presented backwards, “Description is ___,” training the model to retrieve the name. Every single fact is shown in only one direction.
Then comes the real test. For every fact, the model is asked about the direction it has never seen before. The description is presented as a starting point, and the model has to produce the name. The name is presented, and the model has to produce the description. Since the model never trained on the reverse direction, any success there would mean it has somehow generalized the relationship — and any failure would reveal exactly how one-directional the learning is.
## The Results: Perfect One Way, Zero the Other
Training takes only seconds on a standard laptop. The results are stark:
**Trained-direction accuracy: 100%** — the model nails every answer when the question matches the direction it learned.
**Reverse-direction accuracy: 0%** — the model performs no better than random guessing when the question comes from the untrained side.
To put that in perspective, with 400 possible names in the vocabulary, random guessing would yield roughly half a correct answer out of 200 questions. Getting zero is exactly what pure chance looks like. The model learned absolutely nothing usable in the reverse direction.
The gap here mirrors what researchers have observed at much larger scales. The same pattern holds whether you’re working with a model that has 32 numbers per word or 512 — increasing the capacity twelvefold changes nothing. Reversed accuracy stays flat at zero across every size tested.
## Why It Happens: The Mechanical Explanation
The root cause lives in that single line of code that updates only the subject word’s notes during training. Here’s what that means in plain terms:
Every time the model sees a sentence like “Xyphos Brin is the Keeper of Echoes,” it revises its notes on “Xyphos Brin” so it gets better at predicting what comes after that name. But the notes on “Keeper of Echoes” are never touched by that sentence — that phrase was only ever the answer, never the word the prediction started from.
So when the model later encounters “Keeper of Echoes” as if it were the starting point of a question, it reads notes that were never adjusted for that role. They’re still sitting at their random initial values. There’s simply nothing there for the model to work with.
The training process is, in a word, myopic. It optimizes for predicting the next word given what came before, but it doesn’t optimize for the reverse relationship. The connection from A to B gets strengthened; the connection from B to A stays untouched. And since nothing in the training objective rewards forming the reverse link, it never forms.
## What This Means in Practice
The practical implications extend beyond toy models. Any system that relies on a language model’s stored knowledge should be aware that facts aren’t stored symmetrically. Something the model can recall from one phrasing may not come back when you ask from the other direction — and there’s nothing in the output that warns you when that happens.
This is especially relevant for facts that appear infrequently in training data. Common entities and well-known relationships tend to show up in multiple phrasings across the enormous diversity of pretraining text, which may partially mask the problem. But along the long tail of rarer facts, the one-way training scenario is common. A model might know “Elara Voss is the founder of Meridian Labs” perfectly, yet fail when asked “Who founded Meridian Labs?”
It’s also worth contrasting this with a related phenomenon called grokking, where a model gradually discovers a generalizable pattern if it keeps training on a task that rewards finding that pattern. The Reversal Curse is the opposite case: the link that never forms, however long training continues, because nothing in the objective ever asks for it.
## Frequently Asked Questions
**Q: Is this problem fixed by using larger models?**
A: No. The issue persists across model sizes because it’s not about capacity — it’s about the training signal. Even models with vastly more parameters will fail on reverse queries for facts they only learned in one direction, because the reverse connection was never rewarded during training.
**Q: Can the model at least rank the correct answer higher, even if it doesn’t get it right?**
A: Testing reveals no detectable difference. When probed with the reverse question, the model assigns the correct answer roughly the same probability as any other plausible wrong answer. The model isn’t actively avoiding the truth — it simply has no useful signal for it.
**Q: Does this affect all language models equally?**
A: The severity varies. Models pretrained on diverse, massive text corpora encounter facts in many phrasings, which provides some natural coverage of both directions. The problem is most pronounced when a model is trained on a curated dataset where each fact appears predominantly in one direction, or when the fact is rare enough that it only appears in one orientation.
**Q: Can this be worked around with prompting?**
A: Sometimes. If the relevant context is placed in the prompt so the model can reason over it in real time, it can often work out the reverse relationship. The failure mode is specifically about recall from learned weights — retrieving a fact the model memorized during training — not about reasoning with a fact provided in the input.
**Q: Does a simple model like the one described here prove that this is what happens inside large language models?**
A: Not exactly. The toy model demonstrates that one-directional training is sufficient to produce the reversal effect, which is a powerful insight. However, large transformers have many layers, attention mechanisms, and exposure to text where facts appear in multiple orders, so the actual picture inside a real model is more complex. The toy model is a proof of concept, not a complete explanation of what happens at scale.
## Conclusion
The Reversal Curse reveals a fundamental asymmetry in how language models store and retrieve knowledge. They are not symmetric fact databases, and they don’t automatically learn bidirectional relationships from one-directional training examples. The effect holds from toy models with a few dozen numbers per word to large-scale transformers with billions of parameters, suggesting it’s rooted in the training objective itself rather than in architectural limitations.
For anyone building systems that depend on a model’s factual knowledge, this is a critical consideration. Testing knowledge in both directions, not just the direction the model was trained in, can reveal blind spots that would otherwise go undetected. And for researchers, it opens a clear path forward: training objectives and data augmentation strategies that explicitly encourage bidirectional learning could help close this gap.
Thank you for reading



