11 papers
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Mohsen Hariri, Weicong Chen, Nahal Shahini +11
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algori…
Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
Shouren Wang
Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is cap…
MemTrace: Probing What Final Accuracy Misses in Long-Term Memory
Xianxuan Long, Zhikai Chen, Shenglai Zeng +3
LLM agents increasingly maintain long-term memory of user facts across sessions. Yet such memory is usually evaluated by aggregating accuracy over question rows or episodes. Becaus…
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
Wang Yang, Debargha Ganguly, Xinpeng Li +5
Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is lar…
CausalGuard: Conformal Inference under Graph Uncertainty
Vikash Singh, Weicong Chen, Debargha Ganguly +12
Estimating treatment effects from observational data requires choosing an adjustment set, but valid adjustment depends on an unknown causal graph. Graph misspecification can cause…
Reliability-Gated Source Anchoring for Continual Test-Time Adaptation
Vikash Singh, Debargha Ganguly, Weicong Chen +8
Continual test-time adaptation (CTTA) updates a pretrained model online on an unlabeled, non-stationary stream while anchoring it to a frozen source checkpoint. This anchor is usef…