most citedVideo-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2025

DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance

Peiying Zhang, Nanxuan Zhao, Matthew Fisher +3

Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during deco…

cs.CV2025

Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO

Junhao Cheng, Liang Hou, Xin Tao +1

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to d…

cs.GR2025

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

Zhitong Huang, Mohan Zhang, Renhan Wang +3

We present X2Video, the first diffusion model for rendering photorealistic videos guided by intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, whi…

cs.CV2025

Compress Any Segment Anything Model (SAM)

Juntong Fan, Zhiwei Hao, Jianqiang Shen +4

Due to the excellent performance in yielding high-quality, zero-shot segmentation, Segment Anything Model (SAM) and its variants have been widely applied in diverse scenarios such…

cs.CV2025

AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction

Junhao Cheng, Yuying Ge, Yixiao Ge +2

Recent advancements in image and video synthesis have opened up new promise in generative games. One particularly intriguing application is transforming characters from anime films…

cs.CV20251 cited

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Junhao Cheng, Yuying Ge, Teng Wang +3

Recent advances in CoT reasoning and RL post-training have been reported to enhance video reasoning capabilities of MLLMs. This progress naturally raises a question: can these mode…