works on

From the 1 of 25 linked papers with an AI index.

collaborators

25 papers

cs.CV2026

Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting

Xikai Sun, Kebin Liu, Haotian Wang +3

Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs)…

cs.AI2026

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

Haotian Wang, Lian Yan, Xingzhi Yao +4

In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward.…

cs.CV2026

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

Zhuoran Zhang, Bowen Li, Jingcheng Ju +5

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multim…

cs.CV2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Qixun Wang, Yang Shi, Letian Cheng +11

The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…

cs.LG2026

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

Haisen Luo, Yiwei Liu, Haoning Wang +13

Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distilla…

cs.AI2026

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

Yang Shi, Yuhao Dong, Yue Ding +22

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…