4 papers
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models
Yuansheng Gao, Wenbin Xing, Jiahao Yuan +4
Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully suppor…
Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models
Yuansheng Gao, Jinman Zhao, Tong Zhang +5
Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination. Existing inference-time intervention methods u…
VModA: An Effective Framework for Adaptive NSFW Image Moderation
Han Bao, Qinying Wang, Zhi Chen +6
Not Safe/Suitable for Work (NSFW) content is rampant on social networks and poses serious harm to citizens, especially minors. Current detection methods mainly rely on deep learnin…
Boosting Large Language Models for Mental Manipulation Detection via Data Augmentation and Distillation
Yuansheng Gao, Peng Gao, Han Bao +4
Mental manipulation on social media poses a covert yet serious threat to individuals' psychological well-being and the integrity of online interactions. Detecting such behavior is…