activity
20242026
collaborators

5 papers

cs.CV2026

MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer

Zenghao Chai, Chen Tang, Yongkang Wong +2

3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing m…

cs.CV2025

Object-Centric Framework for Video Moment Retrieval

Zongyao Li, Yongkang Wong, Satoshi Yamazaki +2

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such…

cs.LG2025

Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling

Hehe Fan, Yi Yang, Mohan Kankanhalli +1

When modeling a given type of data, we consider it to involve two key aspects: 1) identifying relevant elements (e.g., image pixels or textual words) to a central element, as in a…

cs.CR2024

Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM

Yangyang Guo, Ziwei Xu, Xilie Xu +3

This technical report introduces our top-ranked solution that employs two approaches, \ie suffix injection and projected gradient descent (PGD) , to address the TiFA workshop MLLM…

cs.CV2024

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Wei Li, Hehe Fan, Yongkang Wong +2

Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability…