8 papers
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Diandian Zhang, Tingyu Song, Lin Fu +2
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated…
VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding
Lin Fu, Zheyuan Yang, Yang Wang +3
We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reas…
Parameter Efficient Multi-Class Intelligent Scheduling for Multimodal Online Distributed Industrial Anomaly Detection
Heqiang Wang, Weihong Yang, Zheyuan Yang +4
Industrial anomaly detection has attracted significant attention as a fundamental challenge in industrial systems. The rapid advancement of heterogeneous industrial sensors has dri…
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
Zheyuan Yang, Liqiang Shang, Junjie Chen +6
We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,0…
UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations
Yuan Ren, Guile Wu, Runhao Li +5
Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinh…
Table-R1: Inference-Time Scaling for Table Reasoning
Zheyuan Yang, Lyuhao Chen, Arman Cohan +1
In this work, we present the first study to explore inference-time scaling on table reasoning tasks. We develop and evaluate two post-training strategies to enable inference-time s…