9 papers
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
Zhengbo Zhang, Mark He Huang, Zhigang Tu +1
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query fea…
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22
Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OC…
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
Jixuan He, Xueting Li, Chieh Hubert Lin +1
Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understanding, such as viewpoint reason…
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
Fangzhou Lin, Peiran Li, Lingyu Xu +12
Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture th…
Evolution of Video Generative Foundations
Teng Hu, Jiangning Zhang, Hongrui Huang +7
The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora…
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
Jinxiu Liu, Shaoheng Lin, Yinxiao Li +1
The increasing demand for immersive AR/VR applications and spatial intelligence has heightened the need to generate high-quality scene-level and 360 panoramic video. However, m…