NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
2026-09-30 · Hugging Face
NVIDIA Kumo Tabular: A New Frontier for Tabular Prediction
NVIDIA has introduced Kumo Tabular, an open foundation model for tabular data now available on Hugging Face. As part of the NVIDIA Kumo Structured model collection, it predicts the labels of new rows in a single forward pass, requiring no training, tuning, or feature engineering for both classification and regression tasks.
The Shift to Tabular Foundation Models
Tabular data is the backbone of enterprise machine learning, encompassing customer records, transactions, and sensor logs. For two decades, gradient-boosted trees have dominated this space, but their lifecycle remains rigid—every new task requires collecting labels, engineering features, and training from scratch. Inspired by the in-context learning capabilities of Large Language Models, Kumo Tabular is pretrained on millions of tables. It reads a labeled table as its context and directly predicts the labels of new query rows.
How Kumo Tabular Works
Kumo Tabular is a Transformer built around the structure of a table, utilizing column, row, and in-context attention. To predict labels, it performs three key functions:
- Cell Embedding: A group of cells becomes a token. Numerical and categorical values pass through Fourier features with separate weights. Missing values are treated specially without imputation. Every token in the context receives a label embedding.
- Row Embedding: The model generates row embeddings by alternating column and row attention. Column attention looks down a single column to learn what a value means in its distribution, scaling linearly with the number of rows. Row attention looks across tokens to learn feature interactions using rotary positions. Four learnable [CLS] tokens act as the final readout, making subsequent stages independent of the number of columns.
- In-context Learning: A final Transformer operates on the row embeddings. Context rows attend to each other, while query rows attend only to context rows. Query rows utilize Test-GQA to shrink the read cache. The model outputs class probabilities for classification and 999 quantiles for regression, providing both point predictions and uncertainty estimates.
- Length-aware Attention Temperature: Softmax attention tends to dissolve as the number of keys grows. Kumo Tabular scales every query by a temperature that grows with the logarithm of the number of keys, ensuring attention remains sharp even as tables grow significantly larger or wider.
How Kumo Tabular Was Built
Kumo Tabular is pretrained entirely on artificial tables generated from Structural Causal Models (SCMs). The generation process involves:
1. Drawing a configuration for the whole table, including size, task, mechanisms, and missingness.
2. Creating a random causal graph linking hidden variables, evaluated from root to leaf via randomly drawn functions such as linear maps, small neural networks, trees, or Gaussian processes.
3. Assigning some nodes as numerical or categorical columns, one as the target, and keeping the rest hidden to mimic unmeasured causes in real data.
Released under the OpenMDW-1.1 license for commercial use, Kumo Tabular comes in three sizes ranging from 28M to 215M parameters and ranks first on four major benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench.