5 papers
PhysInOne: Visual Physics Learning and Reasoning in One Suite
Siyuan Zhou, Hejun Wang, Hu Cheng +36
We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to mere…
WebWorld: A Large-Scale World Model for Web Agent Training
Zikai Xiao, Jianhong Tu, Chuhang Zou +7
Web agents require massive trajectories to generalize, yet real-world training is constrained by network latency, rate limits, and safety risks. We introduce \textbf{WebWorld} seri…
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
Xiaodan Hu, Chuhang Zou, Suchen Wang +2
Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models con…
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
Xianghui Xie, Chuhang Zou, Meher Gitika Karumuri +2
We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (us…
Direct and Explicit 3D Generation from a Single Image
Haoyu Wu, Meher Gitika Karumuri, Chuhang Zou +4
Current image-to-3D approaches suffer from high computational costs and lack scalability for high-resolution outputs. In contrast, we introduce a novel framework to directly genera…