14 papers
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Dong Bok Lee, Seanie Lee, Sangwoo Park +12
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning…
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
Sijia Li, Yuchen Huang, Zifan Liu +6
As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing appro…
MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Clinical Tabular Prediction
Zizheng Zhang, Yiming Li, Justin Xu +6
In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasingly used to automate this process, act…
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
Qi Chen, Shuhan Ding, Yu Gu +5
Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, training CT-specific VAEs from scr…
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
Sijia Li, Yuchen Huang, Zifan Liu +7
Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that provide only coarse supervision.…
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
Yuanfang Peng, Jingjing Fu, Chuheng Zhang +6
Reinforcement learning (RL) fine-tuning has shown promise for Vision-Language-Action (VLA) models in robotic manipulation, but deployment-time visual shifts pose practical challeng…