41 papers
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
Chunyang Jiang, Pingping Zhang, Yuzhi Zhao +9
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimod…
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Yiyang Cai, Nan Chen, Rongchang Xie +8
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most app…
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Xinyu Liu, Shihao Li, Weihong Lin +10
The paper introduces ReBind, a framework that uses structured instructions with explicit reference tokens to improve multi‑reference image‑conditioned video editing, enabling preci…
TTHE: Test-Time Harness Evolution
Jun Nie, Yonggang Zhang, Jun Song +5
The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies interm…
AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation
Zeyue Tian, Lei Ke, Zhaoyang Liu +8
Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework,…
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Peng Li, Rawal Khirodkar, Junxuan Li +6
Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressi…