most citedRetrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

6 citations · 8 across the 14 of their papers we have counts for

collaborators

23 papers

cs.CV2025

DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation

Hongfei Zhang, Kanghao Chen, Zixin Zhang +6

This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlled video generation. Recent works have advanced this field by representing camera poses as ra…

cs.CV2025

T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection

Jiazhou Zhou, Qing Jiang, Kanghao Chen +4

Object detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive relian…

cs.CV2025

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

Xu Zheng, Zihao Dongfang, Lutao Jiang +17

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend…

cs.AI2025

AI for Service: Proactive Assistance with AI Glasses

Zichen Wen, Yiyu Wang, Chenfei Liao +10

In an era where AI is evolving from a passive tool into an active and adaptive companion, we introduce AI for Service (AI4Service), a new paradigm that enables proactive and real-t…

cs.CV2025

PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

Zixin Zhang, Kanghao Chen, Xingwang Lin +6

The ability to use, understand, and create tools is a hallmark of human intelligence, enabling sophisticated interaction with the physical world. For any general-purpose intelligen…

cs.CV2025

Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention

Xin Zou, Di Lu, Yizhou Wang +5

Despite their powerful capabilities, Multimodal Large Language Models (MLLMs) suffer from considerable computational overhead due to their reliance on massive visual tokens. Recent…