Study Finds High Dimensionality, Not Formatting, Is Why LLMs Fail at Tabular Prediction
A new research paper investigates why large language models underperform on tabular prediction tasks compared to classical machine learning methods like gradient-boosted trees. Testing nine methods across 31 benchmark datasets, the authors found that LLM accuracy degrades as the number of input features grows, while traditional baselines remain stable or improve. The study ruled out several common explanations — including poor CSV serialization, bad numeric tokenization, and prompt length — as primary causes of this performance gap. Instead, the researchers identified input dimensionality as the central failure mode, arguing that LLMs struggle to detect useful patterns when many weakly related features accumulate. The findings suggest that reformatting data or tweaking prompts is unlikely to resolve the core limitation in high-dimensional tabular learning scenarios.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in