16 papers
Offline-Online Curriculum RL for Multimodal Reasoning
Wendi Deng, Hang Du, Guoshun Nan +11
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines…
SelPE: Progressive Selection for Private Structured Text Synthesis
Xuancheng Zhu, Guoshun Nan, Han Zhang +6
Many data-driven applications rely on structured textual records, such as clinical triage notes and financial transaction logs, for downstream learning and decision-making. In priv…
AutoRAS: Learning Robust Agentic Systems with Primitive Representations
Yang Yue, Xuancheng Zhu, Yuyang Ma +7
The automated design of agentic systems offers a promising pathway for scaling large language models (LLMs) beyond single-agent reasoning. While prior work has advanced task perfor…
See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL
Yilian Liu, Sicong Leng, Guoshun Nan +7
Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffect…
Can LLM Agents Sustain Long-Horizon Organizational Dynamics?
Xuancheng Zhu, Yang Yue, Shuaibing Wan +4
Large language agents are increasingly used for social simulation, yet it remains unclear whether they can sustain coherent behavior in structured organizations, where goals must p…
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
Haolang Lu, Bolun Chu, WeiYe Fu +7
Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between percept…