1 citations · 1 across the 12 of their papers we have counts for
5 papers · 1 filter
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Kaixiang Yao, Xu Wang, Miao Pan +7
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local stat…
Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization
Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang +3
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit…
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Qinfeng Li, Yuntai Bao, Xinyan Yu +7
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, co…
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
Wangjie Gan, Miao Pan, Linbo Xi +4
Large language models are typically post-trained using supervised fine-tuning (SFT) and reinforcement learning (RL), yet effectively unifying efficient knowledge injection with rob…
Compositional Machine Design as Program Synthesis with LLMs
Wenqian Zhang, Yangyi Huang, Weiyang Liu +1
Large language models (LLMs) have shown strong abilities in writing and revising programs, yet many program-synthesis benchmarks still evaluate programs in symbolic or digital envi…