7 papers
Hallucination of Multimodal Large Language Models: A Survey
Zechen Bai, Pichao Wang, Tianjun Xiao +4
This survey presents a comprehensive analysis of the phenomenon of hallucination in multimodal large language models (MLLMs), also known as Large Vision-Language Models (LVLMs), wh…
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
Zechen Bai, Tianjun Xiao, Tong He +4
As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and…
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
Jiaxin Cheng, Zixu Zhao, Tong He +3
Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion an…
Factorized Visual Tokenization and Generation
Zechen Bai, Jianxiong Gao, Ziteng Gao +4
Visual tokenizers are fundamental to image generation. They convert visual data into discrete tokens, enabling transformer-based models to excel at image generation. Despite their…
Unified Lexical Representation for Interpretable Visual-Language Alignment
Yifan Li, Yikai Wang, Yanwei Fu +3
Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clari…
VideoSAM: Open-World Video Segmentation
Pinxue Guo, Zixu Zhao, Jianxiong Gao +5
Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video f…