23 papers
The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling
Shubh Chapra, Dhruv Kumar, Murari Mandal +1
We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the…
FinBalance: A Multi-Document Accounting Reconciliation Benchmark
Sasank Tumpati, Devansh Agarwal, Ayush Kedia +6
Existing financial-NLP benchmarks mostly evaluate prepared artifacts such as filings, tables, or extracted values. Real accounting begins earlier: source documents must be reconcil…
Where Computation Lives Inside TabPFN: Causal Localisation of Attention Head Function
Atharva Gupta, Dhruv Kumar, Murari Mandal +1
We present the first causal mechanistic analysis of a tabular foundation model, investigating how TabPFN 2.5's feature wise attention heads distribute computation across layers. Us…
Mix, Don't Pick: Why Synthetic Corpus Composition Matters for Time Series Foundation Model Pretraining
Aaryan Nagpal, Debdeep Sanyal, Murari Mandal +2
Choosing the wrong synthetic generator for time-series foundation model pretraining is costly: under identical training budgets, the best and worst generators produce up to a $2\ti…
GITCO: Gated Inference-Time Context Optimization in TSFMs
Manya Pandey, Dhruv Kumar, Murari Mandal +1
Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and silently degrade zero-shot fo…
REGEN: Reference-Guided Synthetic Multivariate Time Series Generation for Forecasting
Moulik Gupta, Dhruv Kumar, Murari Mandal +1
Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences. Existing gene…