39 citations · 126 across the 47 of their papers we have counts for
Showing 2025 · cs.CVShow all
3 papers · 2 filters
cs.CV2025
UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
Siqi Li, Xinyu Cai, Jianbiao Mei +7
Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexpl…
cs.CV2025
SymDrive: Realistic and Controllable Driving Simulator via Symmetric Auto-regressive Online Restoration
Zhiyuan Liu, Daocheng Fu, Pinlong Cai +5
High-fidelity and controllable 3D simulation is essential for addressing the long-tail data scarcity in Autonomous Driving (AD), yet existing methods struggle to simultaneously ach…
cs.CV2025
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
Junming Liu, Siyuan Meng, Yanting Gao +7
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially…