activity
20242026
collaborators

29 papers

cs.AI2026

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

Shubh Chapra, Dhruv Kumar, Murari Mandal +1

We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the…

cs.CL2026

FinBalance: A Multi-Document Accounting Reconciliation Benchmark

Sasank Tumpati, Devansh Agarwal, Ayush Kedia +6

Existing financial-NLP benchmarks mostly evaluate prepared artifacts such as filings, tables, or extracted values. Real accounting begins earlier: source documents must be reconcil…

cs.LG2026

Where Computation Lives Inside TabPFN: Causal Localisation of Attention Head Function

Atharva Gupta, Dhruv Kumar, Murari Mandal +1

We present the first causal mechanistic analysis of a tabular foundation model, investigating how TabPFN 2.5's feature wise attention heads distribute computation across layers. Us…

cs.LG2026

Mix, Don't Pick: Why Synthetic Corpus Composition Matters for Time Series Foundation Model Pretraining

Aaryan Nagpal, Debdeep Sanyal, Murari Mandal +2

Choosing the wrong synthetic generator for time-series foundation model pretraining is costly: under identical training budgets, the best and worst generators produce up to a $2\ti…

cs.AI2026

GITCO: Gated Inference-Time Context Optimization in TSFMs

Manya Pandey, Dhruv Kumar, Murari Mandal +1

Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and silently degrade zero-shot fo…

cs.LG2026

REGEN: Reference-Guided Synthetic Multivariate Time Series Generation for Forecasting

Moulik Gupta, Dhruv Kumar, Murari Mandal +1

Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences. Existing gene…