5 papers
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
Rujie Wu, Haozhe Zhao, Hai Ci +1
Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-…
Efficient Action Counting with Dynamic Queries
Xiaoxuan Ma, Zishi Li, Qiuyan Shang +4
Temporal repetition counting aims to quantify the repeated action cycles within a video. The majority of existing methods rely on the similarity correlation matrix to characterize…
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
Fangwei Zhong, Kui Wu, Churan Wang +4
We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environm…
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
Hang Ye, Xiaoxuan Ma, Hai Ci +2
Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skin…
LongViTU: Instruction Tuning for Long-Form Video Understanding
Rujie Wu, Xiaojian Ma, Hai Ci +5
This paper introduces LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. We propose a systematic approach th…