activity
20242026
collaborators

5 papers

cs.CV2026

PhysInOne: Visual Physics Learning and Reasoning in One Suite

Siyuan Zhou, Hejun Wang, Hu Cheng +36

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to mere…

cs.AI2026

WebWorld: A Large-Scale World Model for Web Agent Training

Zikai Xiao, Jianhong Tu, Chuhang Zou +7

Web agents require massive trajectories to generalize, yet real-world training is constrained by network latency, rate limits, and safety risks. We introduce \textbf{WebWorld} seri…

cs.CV2025

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

Xiaodan Hu, Chuhang Zou, Suchen Wang +2

Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models con…

cs.GR2025

MVGBench: Comprehensive Benchmark for Multi-view Generation Models

Xianghui Xie, Chuhang Zou, Meher Gitika Karumuri +2

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (us…

cs.CV2024

Direct and Explicit 3D Generation from a Single Image

Haoyu Wu, Meher Gitika Karumuri, Chuhang Zou +4

Current image-to-3D approaches suffer from high computational costs and lack scalability for high-resolution outputs. In contrast, we introduce a novel framework to directly genera…