activity
20182026
most citedLook Closer to Ground Better: Weakly-Supervised Temporal Grounding of Sentence in Video

49 citations · 111 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

22 papers · 1 filter

cs.CV2026

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Andong Deng, Dawei Du, Zhenfang Chen +7

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…

cs.CV2024

Compositional Physical Reasoning of Objects and Events from Videos

Zhenfang Chen, Shilong Dong, Kexin Yi +5

Understanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and sha…

cs.CV2024

FlexAttention for Efficient High-Resolution Vision-Language Models

Junyan Li, Delin Chen, Tianle Cai +5

Current high-resolution vision-language models encode images as high-resolution image tokens and exhaustively take all these tokens to compute attention, which significantly increa…

cs.CV2024

SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge

Andong Wang, Bo Wu, Sunli Chen +5

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks…

cs.CV2024

ContPhy: Continuum Physical Concept Learning and Reasoning from Videos

Zhicheng Zheng, Xin Yan, Zhenfang Chen +4

We introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense. ContPhy complements existing physical reasoning benchmarks by e…

cs.CV20231 cited

GENOME: GenerativE Neuro-symbOlic visual reasoning by growing and reusing ModulEs

Zhenfang Chen, Rui Sun, Wenjun Liu +2

Recent works have shown that Large Language Models (LLMs) could empower traditional neuro-symbolic models via programming capabilities to translate language into module description…