14 papers
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Dong Bok Lee, Seanie Lee, Sangwoo Park +12
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning…
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Yang Tian, Rui Wang, Xumeng Wen +5
Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide lim…
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
Sijia Li, Yuchen Huang, Zifan Liu +6
As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing appro…
MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Clinical Tabular Prediction
Zizheng Zhang, Yiming Li, Justin Xu +6
In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasingly used to automate this process, act…
In-Context Compositional Q-Learning for Offline Reinforcement Learning
Qiushui Xu, Yuhao Huang, Yushu Jiang +4
Accurate estimation of the Q-function is a central challenge in offline reinforcement learning. However, existing approaches often rely on a shared global Q-function, which is inad…
3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis
Ziyue Wang, Linghan Cai, Chang Han Low +8
3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling…