# A New Era in Retrieval-Augmented Generation: Contextual Embeddings That Understand the Full Document
Retrieval-augmented generation (RAG) systems have transformed how large language models access external knowledge. But a persistent challenge has long undermined their reliability: the way documents are broken into chunks often severs the very connections that make an answer verifiable. A new contextual embedding model aims to solve this problem at the training level, and the results are drawing significant attention across the AI community.
## The Problem With Isolated Chunks
When a long document is split into smaller segments for embedding, each chunk can lose critical context. A sentence like “Monthly rent is $4,500” might seem irrelevant when viewed in isolation, but becomes essential when paired with the lease term, property address, and tenant name that appear elsewhere in the same document. Traditional retrieval models treat each chunk independently, meaning they often retrieve the right paragraph but miss the surrounding context needed to verify whether the answer is actually correct.
This is where the concept of **late chunking** comes in. Instead of embedding chunks in isolation, late chunking encodes the entire document in a single pass and then pools representations per chunk. The result is that every chunk carries a summary of the full document, making retrieval far more robust to cross-sentence dependencies.
However, late chunking alone does not fix how the model is trained. Most training pipelines label exactly one “gold” chunk per query and treat every other chunk as a negative — including chunks that contain the supporting evidence needed to verify the answer. This coarse, binary signal limits what the model can learn.
## A Token-Level Teacher Replaces the Gold Passage
The breakthrough lies in a new training signal. Rather than designating a single correct chunk, the model uses a **teacher model** that reads the query and the full document together, scoring every token for relevance. Those token-level scores are then aggregated per chunk — specifically, the mean of the top-n token scores within each chunk.
This produces a **soft target**: a probability distribution over all chunks in the positive document, rather than a single one-hot label. Chunks that contain even partial evidence for the answer receive a graded signal proportional to their relevance. A **distillation loss**, computed as forward KL divergence between the teacher’s distribution and the student’s distribution, guides the student model to match these nuanced soft targets.
Additionally, a **document-level loss** inspired by ColBERT’s MaxSim mechanism ensures that each document is scored by its most relevant chunk. Each training batch also samples a random chunking strategy, meaning the model learns boundaries that are flexible rather than fixed to a single segmentation scheme.
Crucially, the teacher model is only used during training. At inference time, the student model operates independently with no added latency or storage overhead.
## Model Architecture and Training Details
The student model builds on an in-house 9-billion-parameter ColBERT-style retrieval architecture (reported as 8B on Hugging Face). A linear projection layer maps representations to a 2,048-dimensional embedding space, with Matryoshka training enabling 1,024-dimensional outputs as well. Quantization-aware training during the pre-training phase allows the model to produce native int8 embeddings without any post-training conversion.
The release is described as a model soup, combining several checkpoints. Training drew from approximately 430 datasets spanning over 50 languages, though notably excluded data from the ConTEB benchmark to prevent contamination.
Weights are available on Hugging Face under the MIT license. Loading requires `transformers` version 5.4.0 or higher with `trust_remote_code=True`. The model is currently a self-hosted preview and has not yet been made available through the Perplexity API. The model card cautions that weights and the interface may change without backward compatibility.
## context-bench: A Purpose-Built Benchmark
To evaluate the model’s contextual retrieval capabilities, a new benchmark called **context-bench** was introduced. It is privately held by turbopuffer to prevent training contamination. The benchmark consists of 2,099 queries across 38,894 documents in 21 domains, with a median target document length of roughly 6,100 tokens.
Using sentence-level chunking, the dataset yields over 2.4 million chunks. The queries are designed to test 12 distinct contextual capabilities, ranging from pronoun resolution and coreference to understanding table structures and multi-sentence reasoning.
Evaluation uses four metrics: **Document@K**, **Answer@K**, **Evidence Recall@K**, and **All-Evidence@K**. Every model is ranked exhaustively against all chunks, which means index settings, batch sizes, and retrieval parameters have no effect on the reported numbers.
## Reported Results
At K = 10, the model achieves **45.5% answer recall** and **40.6% evidence recall**, with **31.1% all-evidence recall**. Document-level recall stands at 15.2% at K = 1 and climbs to 61.6% at K = 10.
On the **context-bench leaderboard**, the model outperforms the closest contextual competitor by a substantial margin: **14.4 points** on answer recall and **5.0 points** on evidence recall at K = 10. Among general retrieval benchmarks, it holds the highest average nDCG@10 on the ConTEB suite and leads on query-to-chunk tasks, while running slightly behind the competitor on query-to-document tasks.
Perhaps most striking is the storage efficiency. The 1,024-dimensional int8 variant — storing just **1 KB per vector** — slightly outperforms the competitor’s 2,048-dimensional float32 version (8 KB per vector) on the chunk-retrieval suite. This suggests that contextual embeddings do not require large-dimensional dense vectors to achieve state-of-the-art performance.
## Storage and Cost Considerations
Vector storage scales linearly with the number of chunks, dimensions, and bytes per value. At scale, this difference is substantial. For a corpus of 2.5 million chunks:
– **1,024-dim int8** (1 KB/vector): roughly 2.5 GB
– **2,048-dim float32** (8 KB/vector): roughly 20 GB
For enterprise deployments managing hundreds of millions of documents, this order-of-magnitude reduction in storage requirements translates directly into lower infrastructure costs and faster indexing speeds.
The model’s chunk-size sensitivity is also noteworthy. Across 74 MTEB tasks, the mean nDCG@10 moves only from 81.0% at 64 tokens per chunk to 79.9% at 512 tokens per chunk, indicating robustness across a wide range of chunk granularities.
## Comparison With Existing Models
The model distinguishes itself on several axes when compared to contemporary embedding systems. It is the only open-weight model in its class to offer native int8 quantized output with a contextual training paradigm. While competitors rely on hosted APIs for contextual embeddings, this model can be run entirely on-premises, giving organizations full control over their data and retrieval pipelines.
The comparison table below summarizes the key differentiators across leading models in the space:
| Feature | This Model | Leading Contextual API Model | Prior Open-Weight Contextual | Independent Chunk Model |
|—|—|—|—|—|
| **Chunk Embeddings** | Contextual | Contextual | Contextual | Independent per chunk |
| **License** | MIT | Proprietary API | MIT | OpenMDW-1.1 |
| **Parameters** | ~9B (8B on HF) | Not disclosed | 4B | ~8B |
| **Dimensions** | 2048, 1024 | Multiple, including 256 | 2560 (Matryoshka) | 4096, sliceable |
| **Quantized Output** | Native int8 | int8, uint8, binary | int8, binary | Float only |
| **Max Context** | 32,768 tokens | 32K (120K with auto-chunking) | 32K | 32,768 |
| **Auto-Chunking** | No | Yes | No | No |
| **Deployment** | Self-hosted | API-only | Self-host or API | Self-hosted |
## What This Means for RAG Pipelines
The practical implication is straightforward: RAG systems can now retrieve not just a passage that contains the answer, but a passage where the answer is **verifiable** in context. This is especially valuable for domains like legal document retrieval, financial analysis, medical literature search, and technical support, where a sentence extracted from the middle of a document may be misleading without the surrounding clauses or definitions.
The ability to train with flexible chunk boundaries also means that organizations can experiment with different chunking strategies — from short 64-token snippets to longer 512-token passages — without needing to re-annotate their training data or retrain the model from scratch.
## Limitations and Open Questions
The model is currently released as a preview, with no backward compatibility guaranteed for future updates. It is not yet accessible through any managed API, which may limit adoption for teams without the infrastructure to self-host large models. Additionally, while the model performs strongly on contextual retrieval tasks, it is slightly behind the top competitor on pure query-to-document ranking tasks, suggesting trade-offs between contextual depth and broad retrieval coverage.
The exclusion of ConTEB data from training raises questions about how the model would perform on that benchmark if fine-tuned with it, though the authors deliberately avoided this to maintain benchmark integrity.
## FAQ
**What makes this embedding model different from standard embedding models?**
It uses a token-level teacher model that scores every token for relevance during training, producing soft targets instead of a single gold chunk label. This allows the model to learn nuanced contextual relationships across entire documents.
**Can I use this model in production today?**
Yes, but only in a self-hosted capacity. The weights are available on Hugging Face under the MIT license, and you will need `transformers>=5.4.0` to load them. It is not yet available through any hosted API.
**How does late chunking work?**
The entire document is encoded in a single forward pass, and representations are pooled per chunk afterward. This ensures every chunk inherits information from the full document, not just the sentences it contains.
**What is context-bench?**
A privately held benchmark with 2,099 queries over 38,894 documents in 21 domains. It was specifically designed to evaluate contextual retrieval capabilities and to prevent training contamination.
**What does the 1 KB vector size mean for storage costs?**
At 1,024 dimensions with int8 quantization, each vector uses 1 KB. For 1 million chunks, this is approximately 1 GB of storage, compared to roughly 8 GB for a 2,048-dim float32 model.
**Does the model support auto-chunking?**
No. The model does not include an automatic chunking mechanism. Users need to define their own chunking strategy before embedding documents.
**What languages does the model support?**
Training covered over 50 languages across approximately 430 datasets. Specific language performance is detailed in the model card on Hugging Face.
**What chunk size is recommended?**
The model is robust across chunk sizes from 64 to 512 tokens, with minimal performance degradation. The optimal choice depends on the use case and the typical document structure in your corpus.
## Conclusion
The release of this contextual embedding model represents a meaningful step forward for retrieval-augmented generation. By replacing the single gold-chunk paradigm with token-level teacher distillation, it captures the nuance of document-level context at scale. The combination of strong benchmark performance, MIT-licensed open weights, and remarkable storage efficiency makes it a compelling option for teams building production RAG systems.
While challenges remain — including the lack of API access and the preview status of the weights — the direction is clear: future embedding models will need to understand not just individual passages, but the full reasoning chain that makes an answer trustworthy. This model lays important groundwork for that future.
Thank you for reading



