8 papers
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
Jiakang Li, Guanyu Zhu, Can Jin +8
Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavio…
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
Dexu Yu, Youhua Li, Zhaoyang Guan +12
Agent skills have become a practical way to extend large language model agents, but the growing skill ecosystem still lacks a reliable way to judge whether a skill is worth deployi…
GIFT: LLM-Guided State-Reward Interface for Financial Reinforcement Learning
Yanyan Wu, Boyi Zhang, Yanlin Liu +10
Financial portfolio trading is naturally formulated as a reinforcement learning problem, where an agent sequentially rebalances assets under changing market conditions to balance r…
On the Role of Language Representations in Auto-Bidding: Findings and Implications
Guanyu Zhu, Jining Luan, Hanwen Du +11
Auto-bidding is a crucial task in real-time advertising markets, where policies must optimize long-horizon value under delivery constraints (e.g., budget and CPA). Existing methods…
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…
CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation
Junchen Fu, Yongxin Ni, Joemon M. Jose +4
In this paper, we explore a less-studied yet practically important problem: how to efficiently and effectively adapt multiple (2) multimodal foundation models (MFMs) for the seq…