activity
20242026
collaborators

5 papers

cs.CV2026

Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation

Yanliang Qi, Kexi Chen, Muchao Ye +1

Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must simultaneously preserve the ap…

cs.CV2026

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

Xin Hu, Haomiao Ni, Yunbei Zhang +3

Vision language models (VLMs) have achieved remarkable success in broad visual understanding, yet they remain challenged by object-centric reasoning on rare objects due to the scar…

cs.CV2025

SafeTriage: Facial Video De-identification for Privacy-Preserving Stroke Triage

Tongan Cai, Haomiao Ni, Wenchao Ma +8

Effective stroke triage in emergency settings often relies on clinicians' ability to identify subtle abnormalities in facial muscle coordination. While recent AI models have shown…

cs.CV2025

Computer-Aided Layout Generation for Building Design: A Review

Jiachen Liu, Yuan Xue, Haomiao Ni +3

Generating realistic building layouts for automatic building design has been studied in both the computer vision and architecture domains. Traditional approaches from the architect…

cs.CV2024

TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models

Haomiao Ni, Bernhard Egger, Suhas Lohit +5

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., "a woman is…