Publications (5)
A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models
Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng +2
Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graph…
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
Xuanming Cui, Hong-You Chen, Hao Yu +10
Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (T…
Think Then Embed: Generative Context Improves Multimodal Embedding
Xuanming Cui, Jianpeng Cheng, Hong-you Chen +11
There is a growing interest in Universal Multimodal Embeddings (UME), where models are required to generate task-specific representations. While recent studies show that Multimodal…
AirSketch: Generative Motion to Sketch
Hui Xian Grace Lim, Xuanming Cui, Yogesh S Rawat +1
Illustration is a fundamental mode of human expression and communication. Certain types of motion that accompany speech can provide this illustrative mode of communication. While A…
On the Robustness of Large Multimodal Models Against Image Adversarial Attacks
Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang +1
Recent advances in instruction tuning have led to the development of State-of-the-Art Large Multimodal Models (LMMs). Given the novelty of these models, the impact of visual advers…