1 citations · 1 across the 4 of their papers we have counts for
7 papers
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
Peiying Zhang, Nanxuan Zhao, Matthew Fisher +3
Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during deco…
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
Junhao Cheng, Liang Hou, Xin Tao +1
While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to d…
X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
Zhitong Huang, Mohan Zhang, Renhan Wang +3
We present X2Video, the first diffusion model for rendering photorealistic videos guided by intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, whi…
Compress Any Segment Anything Model (SAM)
Juntong Fan, Zhiwei Hao, Jianqiang Shen +4
Due to the excellent performance in yielding high-quality, zero-shot segmentation, Segment Anything Model (SAM) and its variants have been widely applied in diverse scenarios such…
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
Junhao Cheng, Yuying Ge, Yixiao Ge +2
Recent advancements in image and video synthesis have opened up new promise in generative games. One particularly intriguing application is transforming characters from anime films…
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
Junhao Cheng, Yuying Ge, Teng Wang +3
Recent advances in CoT reasoning and RL post-training have been reported to enhance video reasoning capabilities of MLLMs. This progress naturally raises a question: can these mode…