16 papers
Improving Generalization Robustness of Multimodal RLVR
Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng +11
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing th…
Unified Hallucination Fuzzing for Multimodal Large Language Models
Pengfei Zhou, Jiajun Song, Zhiwei Tang +12
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, pr…
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs
Weidong Tang, Jierui Li, Yueling Hou +7
True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental states interact and crystallize…
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
Xiaoda Yang, Majun Zhang, Changhao Pan +10
Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from g…
On-the-Fly Data Augmentation via Gradient-Guided and Sample-Aware Influence Estimation
Suorong Yang, Jie Zong, Lihang Wang +6
Data augmentation has been widely employed to improve the generalization of deep neural networks. Most existing methods apply fixed or random transformations. However, we find that…
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
Wangbo Zhao, Yizeng Han, Zhiwei Tang +7
Diffusion Transformers (DiTs) excel at visual generation yet remain hampered by slow sampling. Existing training-free accelerators - step reduction, feature caching, and sparse att…