5 papers
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
Shuai Wang, Liang Li, Yang Chen +3
Unified Multimodal Models (UMMs) have emerged as a critical direction for general-purpose multimodal intelligence, integrating understanding and generation into a single framework.…
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
Chunxu Liu, Jiyuan Yang, Ruopeng Gao +4
Multimodal embeddings are widely used in downstream tasks such as multimodal retrieval, enabling alignment of interleaved modalities in a shared representation space. While recent…
SORCE: Small Object Retrieval in Complex Environments
Chunxu Liu, Chi Xie, Xiaxu Chen +4
Text-to-Image Retrieval (T2IR) is a highly valuable task that aims to match a given textual query to images in a gallery. Existing benchmarks primarily focus on textual queries des…
Multiple Object Tracking as ID Prediction
Ruopeng Gao, Ji Qi, Limin Wang
Multi-Object Tracking (MOT) has been a long-standing challenge in video understanding. A natural and intuitive approach is to split this task into two parts: object detection and a…
History-Aware Transformation of ReID Features for Multiple Object Tracking
Ruopeng Gao, Yuyao Wang, Chunxu Liu +1
In Multiple Object Tracking (MOT), Re-identification (ReID) features are widely employed as a powerful cue for object association. However, they are often wielded as a one-size-fit…