4 papers
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
Lihong Wang, Liangqi Li, Weiwei Feng +6
CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks. Existing MLLMs typic…
Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
Lin Chen, Xiaoke Zhao, Kun Ding +9
Multimodal Large Language Models (MLLMs) demonstrate impressive cross-modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation…
DDL: A Large-Scale Datasets for Deepfake Detection and Localization in Diversified Real-World Scenarios
Changtao Miao, Yi Zhang, Weize Gao +11
Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this…
MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
Changtao Miao, Yi Zhang, Man Luo +9
Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However,…