3 papers
cs.CL2026
Yunque DeepResearch Technical Report
Yuxuan Cai, Xinyi Lai, Peng Yuan +8
Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full…
cs.CV2025
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
Zhaoyang Wei, Wenchao Ding, Yanchao Hao +1
Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing thi…
cs.CV2024
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
Zhengfei Xu, Sijia Zhao, Yanchao Hao +6
Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Pre…