47 citations · 121 across the 23 of their papers we have counts for
30 papers · 1 filter
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Chongjian Ge, Hanwen Jiang, Tianyu Wang +9
Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera,…
SimpleMatch: A Simple and Strong Baseline for Semantic Correspondence
Hailing Jin, Huiying Li
Recent advances in semantic correspondence have been largely driven by the use of pre-trained large-scale models. However, a limitation of these approaches is their dependence on h…
Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training
Dezhao Luo, Jiabo Huang, Shaogang Gong +2
The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for vi…
LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos
Jielin Qiu, Franck Dernoncourt, Trung Bui +3
Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the s…
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automa…
MHMS: Multimodal Hierarchical Multimedia Summarization
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or p…