1 citations · 3 across the 21 of their papers we have counts for
6 papers · 1 filter
ViMax: Agentic Video Generation
Lingxuan Huang, Sizhe He, Hengji Zhou +3
Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…
DramaDirector: Geometry-Guided Short Drama Generation
Hengji Zhou, Sijie Liu, Jianrun Chen +3
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation…
VideoAgent: All-in-One Framework for Video Understanding and Editing
Hengji Zhou, Lingxuan Huang, Jian Wang +4
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two cri…
Reading Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
Heng Zhou, Ao Yu, Li Kang +5
Vision-Language Models achieve near-perfect accuracy at reading text in images, yet prove largely typography-blind: capable of recognizing what text says, but not how it looks. We…
From Perception to Action: An Interactive Benchmark for Vision Reasoning
Yuhao Wu, Maojia Song, Yihuai Lan +8
Understanding the physical structure is essential for real-world applications such as embodied agents, interactive design, and long-horizon manipulation. Yet, prevailing Vision-Lan…
SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset
Yubin Hu, Kairui Wen, Heng Zhou +2
Reconstructing accurate 3D surfaces for street-view scenarios is crucial for applications such as digital entertainment and autonomous driving simulation. However, existing street-…