6 papers · 1 filter
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
Yanbing Zhang, Bo Wang, Jianhui Liu +9
Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static obse…
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7
Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual questio…
Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations
Yuan Zhang, Chenyi Li, Guoqing Ma +8
Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often…
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Guangcong Zheng, Jianlong Yuan, Bo Wang +3
Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use aut…
STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
Bo Wang, Haoyang Huang, Zhiying Lu +6
This paper introduces StoryAnchors, a unified framework for generating high-quality, multi-scene story frames with strong temporal consistency. The framework employs a bidirectiona…
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
Haoyang Huang, Guoqing Ma, Nan Duan +51
We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and…