activity
20242026
most citedState Rank Dynamics in Linear Attention LLMs

1 citations · 3 across the 21 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

ViMax: Agentic Video Generation

Lingxuan Huang, Sizhe He, Hengji Zhou +3

Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…

cs.CV2026

DramaDirector: Geometry-Guided Short Drama Generation

Hengji Zhou, Sijie Liu, Jianrun Chen +3

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation…

cs.CV2026

VideoAgent: All-in-One Framework for Video Understanding and Editing

Hengji Zhou, Lingxuan Huang, Jian Wang +4

Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two cri…

cs.CV2026

Reading Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models

Heng Zhou, Ao Yu, Li Kang +5

Vision-Language Models achieve near-perfect accuracy at reading text in images, yet prove largely typography-blind: capable of recognizing what text says, but not how it looks. We…

cs.CV2026

From Perception to Action: An Interactive Benchmark for Vision Reasoning

Yuhao Wu, Maojia Song, Yihuai Lan +8

Understanding the physical structure is essential for real-world applications such as embodied agents, interactive design, and long-horizon manipulation. Yet, prevailing Vision-Lan…

cs.CV20241 cited

SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset

Yubin Hu, Kairui Wen, Heng Zhou +2

Reconstructing accurate 3D surfaces for street-view scenarios is crucial for applications such as digital entertainment and autonomous driving simulation. However, existing street-…