collaborators

40 papers

cs.CV2026

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Hongxing Li, Xiufeng Huang, Dingming Li +11

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approach…

cs.CL2026

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

Yuchen Yan, Liang Jiang, Jin Jiang +7

Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, context length limits, and degraded rea…

cs.CV2026

InstructSAM: Segment Any Instance with Any Instructions

Yuqian Yuan, Wentong Li, Zhaocheng Li +6

In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven…

cs.CL2026

GroundAct: Can LLM Agents Ground Actions in Environmental States?

Zixuan Wang, Dingming Li, Hongxing Li +8

LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instructi…

cs.CL2026

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

Haoyu Zheng, Yun Zhu, Shu Yuan +5

Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lost-in-conversation (LiC) gap…

cs.CV2026

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

Wei Wang, Yuqian Yuan, Tianwei Lin +4

Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and intera…