116 citations · 225 across the 27 of their papers we have counts for
27 papers
Controlled Dynamics Attractor Transformer
Cheng Zhang, Minnan Luo, Zesheng Yang +3
Transformer architectures have dramatically advanced representation learning and inference in deep models through self-attention mechanisms. In parallel,associative memory (AM) fra…
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
Haochen Han, Qinghua Zheng, Guang Dai +2
Collecting well-matched multimedia datasets is crucial for training cross-modal retrieval models. However, in real-world scenarios, massive multimodal data are harvested from the I…
MMoE: Robust Spoiler Detection with Multi-modal Information and Domain-aware Mixture-of-Experts
Zinan Zeng, Sen Ye, Zijian Cai +4
Online movie review websites are valuable for information and discussion about movies. However, the massive spoiler reviews detract from the movie-watching experience, making spoil…
Disentangled Representation Learning with Transmitted Information Bottleneck
Zhuohang Dang, Minnan Luo, Chengyou Jia +4
Encoding only the task-related information from the raw data, \ie, disentangled representation learning, can greatly contribute to the robustness and generalizability of models. Al…
PSDiff: Diffusion Model for Person Search with Iterative and Collaborative Refinement
Chengyou Jia, Minnan Luo, Zhuohang Dang +3
Dominant Person Search methods aim to localize and recognize query persons in a unified network, which jointly optimizes two sub-tasks, \ie, pedestrian detection and Re-IDentificat…
Clustering based Point Cloud Representation Learning for 3D Analysis
Tuo Feng, Wenguan Wang, Xiaohan Wang +2
Point cloud analysis (such as 3D segmentation and detection) is a challenging task, because of not only the irregular geometries of many millions of unordered points, but also the…