3 papers
cs.LG2026
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
Gagan Mundada, Zihan Huang, Rohan Surana +8
Group Relative Policy Optimization (GRPO) is effective for training language models on complex reasoning. However, since the objective is defined relative to a group of sampled tra…
cs.AI2026
Insight Agents: An LLM-Based Multi-Agent System for Data Insights
Jincheng Bai, Zhenyu Zhang, Jennifer Zhang +1
Today, E-commerce sellers face several key challenges, including difficulties in discovering and effectively utilizing available programs and tools, and struggling to understand an…
cs.CV2026
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang +5
Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each st…