3 papers
cs.AI2026
Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning
Bincheng Gu, Min Gao, Zongwei Wang +3
Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks.…
cs.LG2026
Test-Time Personalization: A Diagnostic Framework and Probabilistic Fix for Scaling Failures
Linhai Zhang, Yulan He
Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study Tes…
cs.CL2025
Hallucinations are inevitable but can be made statistically negligible
Atsushi Suzuki, Yulan He, Feng Tian +1
Hallucinations, a phenomenon where a language model (LM) generates nonfactual content, pose a significant challenge to the practical deployment of LMs. While many empirical methods…