7 papers
VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
Tianxiao Chen, Hanmo Chen, Huajin Chen +3
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and im…
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
Bo Yin, Xiaobin Hu, Xingyu Zhou +7
Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains challenging. We revisit the reco…
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
Juncheng Ma, Yuxuan Du, Yanan Sun +8
Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although traini…
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
Bo Yin, Xiaobin Hu, Chengming Xu +6
Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evi…
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems
Yunhang Qian, Xiaobin Hu, Jiaquan Yu +6
While Multi-Agent Systems (MAS) show potential for complex clinical decision support, the field remains hindered by architectural fragmentation and the lack of standardized multimo…
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
Guangyuan Li, Bo Li, Jinwei Chen +3
Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First,…