collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

Yanbing Zhang, Bo Wang, Jianhui Liu +9

Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static obse…

cs.CV2026

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7

Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual questio…

cs.CV2026

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

Yuan Zhang, Chenyi Li, Guoqing Ma +8

Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often…

cs.CV2025

Frame-Level Captions for Long Video Generation with Complex Multi Scenes

Guangcong Zheng, Jianlong Yuan, Bo Wang +3

Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use aut…

cs.CV2025

STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives

Bo Wang, Haoyang Huang, Zhiying Lu +6

This paper introduces StoryAnchors, a unified framework for generating high-quality, multi-scene story frames with strong temporal consistency. The framework employs a bidirectiona…

cs.CV2025

Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Haoyang Huang, Guoqing Ma, Nan Duan +51

We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and…