# NVIDIA Unveils Kumo Tabular: A New Family of Foundation Models for Tabular Data
## A Zero-Training Approach to Structured Data Predictions
NVIDIA has introduced Kumo Tabular, a groundbreaking family of tabular foundation models designed for both classification and regression tasks on structured datasets. Unlike traditional machine learning pipelines that require extensive preprocessing, model selection, and iterative tuning, Kumo Tabular operates on a remarkably simple premise: feed it labeled rows and receive predictions for new rows in a single forward pass.
This approach draws from recent advances in tabular foundation modeling, following in the footsteps of prior work that demonstrated Transformer architectures could be effectively applied to table-shaped data. What sets Kumo Tabular apart is its emphasis on ease of deployment and practical usability, making it accessible to practitioners who need fast, reliable predictions without the overhead of a full training loop.
—
## Model Variants and Scale
The Kumo Tabular family spans three model sizes: Small, Medium, and Large. These variants range from approximately 28 million parameters all the way up to roughly 215 million parameters. This range gives users the flexibility to choose the right balance between computational cost and predictive performance depending on their specific use case.
All three variants are available through NVIDIA’s open-source structured-data-models library, commonly referred to as SDM. The library serves as a unified framework for working with tabular foundation models and includes support for multiple model architectures beyond just Kumo Tabular.
—
## Licensing and Deployment
One of the most significant practical advantages of Kumo Tabular is its permissive licensing. The model weights are distributed under the OpenMDW-1.1 license, which explicitly allows commercial use. The underlying SDM library code itself is released under the Apache 2.0 license, further lowering barriers to adoption in production environments.
To run the models, users need Python 3.11 or newer, PyTorch 2.7 or later, and a CUDA-compatible GPU. NVIDIA has provided example workflows that target GPU acceleration out of the box, ensuring that users can get started with minimal infrastructure setup.
—
## What the SDM Library Brings to the Table
The SDM library is more than just a wrapper for Kumo Tabular. It is a GPU-native framework purpose-built for structured data foundation models and preprocessing. Beyond Kumo Tabular, the library includes TabICLv2, Google’s TabFM, and KumoRelational, which is specifically designed for multi-table relational data.
A key design choice across all these models is a shared in-context learning interface built on a unified data container called TableTensor. This standardization means that switching between different models or combining them into ensembles follows a consistent API. The library also handles preprocessing steps automatically and supports ensembling strategies and many-class prediction scenarios.
—
## How Kumo Tabular Works
### The Three-Stage Pipeline
At its core, Kumo Tabular is a Transformer architecture specifically tailored to the structure of tabular data. It draws on established techniques like column attention, row attention, and in-context learning that were pioneered in earlier tabular foundation model research. The processing pipeline unfolds in three distinct stages:
**Stage 1: Cell Embedding**
Each individual cell in the table — whether it holds a numerical value or belongs to a categorical feature — is transformed through learned Fourier feature mappings. Numerical and categorical data types receive separate embedding weights, allowing the model to learn type-specific representations. Notably, missing values are handled natively without requiring any imputation step. The model learns dedicated representations for missingness, eliminating the need for users to preprocess or fill in gaps beforehand.
**Stage 2: Row Embedding**
Once cells are embedded, the model constructs row-level representations using two types of attention. Column attention operates downward through individual columns and uses an induced self-attention mechanism, which keeps computational cost growing linearly rather than quadratically with the number of rows. Row attention then operates across columns within each row, learning interactions between features. Rotary position encodings are applied so the model can distinguish one column from another. After row attention, four learnable special tokens compress each row into a compact summary representation. This means the downstream computational cost becomes independent of the number of columns.
**Stage 3: In-Context Learning**
The final stage runs a Transformer over the row embeddings. Context rows — those with known labels — attend to one another. Query rows — those without labels — attend exclusively to the context rows. Because context rows never see the query rows, the keys and values from context rows can be computed once and then reused for multiple query predictions. The output head produces either class probabilities for classification tasks or 999 quantile estimates for regression tasks. This design gives users not just a single point prediction but also a built-in uncertainty estimate.
### Scaling Attention for Large Tables
A subtle but important engineering detail addresses the challenge of softmax attention diluting as the number of keys increases. Kumo Tabular applies a temperature scaling factor to each query that grows logarithmically with the number of keys in the context. Crucially, this coefficient is learned independently for each attention head during pretraining, allowing the model to maintain sharp attention even when processing tables with tens of thousands of rows.
—
## Pretraining and Capabilities
All Kumo Tabular variants were pretrained exclusively on artificial tables generated from structural causal models. This synthetic pretraining strategy allows the models to learn generalizable patterns about how features relate to one another and to target variables, without being tied to any specific domain.
For classification, the models output probabilities for up to 10 classes in a single forward pass, with support for extending beyond that limit using error-correcting output codes. For regression, the models emit 999 quantile estimates — from the 0.1st percentile to the 99.9th percentile — giving users a full distributional view of their predictions rather than just a single number.
—
## Frequently Asked Questions
**What types of tasks can Kumo Tabular handle?**
Kumo Tabular supports both classification and regression on structured tabular data. It can handle numerical and categorical features, missing values, and many-class classification problems.
**Do I need to train the model myself?**
No. Kumo Tabular is a pretrained foundation model. You provide labeled rows as context and receive predictions for unlabeled rows — no training loop, no gradient computation, and no hyperparameter tuning required.
**What hardware do I need?**
A CUDA-compatible GPU is required. The models are built on PyTorch and target GPU acceleration for efficient inference.
**Can I use Kumo Tabular in a commercial product?**
Yes. The model weights are released under the OpenMDW-1.1 license, which permits commercial use. The SDM library code is Apache-2.0 licensed.
**How does Kumo Tabular handle missing data?**
Missing values are handled natively by the cell embedding stage. The model learns dedicated representations for missingness, so users do not need to impute or drop missing entries.
**What happens if I have more than 10 classes?**
Kumo Tabular supports up to 10 classes per forward pass. For problems with more classes, the model can be extended using error-correcting output codes.
**Can I use multiple model sizes together?**
Yes. The SDM library supports ensembling, and all model sizes share a consistent interface through the TableTensor container, making it straightforward to combine predictions from Small, Medium, and Large variants.
**What is the difference between Kumo Tabular and other tabular foundation models?**
Kumo Tabular distinguishes itself through its temperature-scaled attention mechanism for handling large tables, its built-in uncertainty estimates via quantile regression, and its integration into a comprehensive library that supports multiple model architectures and preprocessing pipelines.
—
## Conclusion
NVIDIA’s Kumo Tabular represents a significant step forward in making foundation-model-style approaches practical for tabular data. By eliminating the need for training, hyperparameter tuning, and manual feature engineering, it dramatically reduces the friction between having a structured dataset and getting actionable predictions. The combination of permissive licensing, a well-documented open-source library, and multiple model sizes ensures that both researchers and industry practitioners can adopt it with confidence. As tabular foundation models continue to mature, Kumo Tabular sets a compelling benchmark for what zero-shot structured data prediction can look like in practice.
Thank you for reading



