collaborators

6 papers

cs.RO2026

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure

Yiyuan Pan, Hanjiang Hu, Shangtao Li +2

A central challenge in deploying learned robot policies is inference-time behavior steering: redirecting a policy at test time to satisfy user preferences not anticipated during tr…

cs.CL2026

PhoneBuddy: Training Open Models for Agentic Phone Use

Zhengyang Tang, Xin Lai, Pengyuan Lyu +23

Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matter…

cs.AI2026

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Haoling Li, Kai Zheng, Jie Wu +4

Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labels themselves must remain reli…

cs.LG2026

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Haipeng Luo, Qingfeng Sun, Songli Wu +4

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from poli…

cs.AI2026

AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent

Haipeng Luo, Huawen Feng, Qingfeng Sun +6

Large Reasoning Models (LRMs) like o3 and DeepSeek-R1 have achieved remarkable progress in reasoning tasks with long cot. However, they remain computationally inefficient and strug…

cs.CV2025

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li +15

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…