9 papers
Object-Centric Framework for Video Moment Retrieval
Zongyao Li, Yongkang Wong, Satoshi Yamazaki +2
Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such…
MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
Zenghao Chai, Chen Tang, Yongkang Wong +2
3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing m…
Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM
Yangyang Guo, Ziwei Xu, Xilie Xu +3
This technical report introduces our top-ranked solution that employs two approaches, \ie suffix injection and projected gradient descent (PGD) , to address the TiFA workshop MLLM…
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
Zenghao Chai, Chen Tang, Yongkang Wong +1
The creation of 4D avatars (i.e., animated 3D avatars) from text description typically uses text-to-image (T2I) diffusion models to synthesize 3D avatars in the canonical space and…
Bridging the Intent Gap: Knowledge-Enhanced Visual Generation
Yi Cheng, Ziwei Xu, Dongyun Lin +5
For visual content generation, discrepancies between user intentions and the generated content have been a longstanding problem. This discrepancy arises from two main factors. Firs…
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
Wei Li, Hehe Fan, Yongkang Wong +2
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability…