2 papers
cs.LG2026
Interpretable Tabular Foundation Models via In-Context Kernel Regression
Ratmir Miftachov, Bruno Charron, Simon Valentin
Tabular foundation models like TabPFN and TabICL achieve state-of-the-art performance through in-context learning, yet their architectures remain fundamentally opaque. We introduce…
cs.SE2025
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10
Coding agents powered by large language models have shown impressive capabilities in software engineering tasks, but evaluating their performance across diverse programming languag…