8 papers
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
Haowen Gao, Zhenyu Zhang, Liang Pang +7
Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…
Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
Jun Li, Hongjian Dou, Zhenyu Zhang +3
Composed Image Retrieval (CIR) presents a significant challenge as it requires jointly understanding a reference image and a modified textual instruction to find relevant target im…
Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
Chenyu Zhou, Xiaoming Shi, Hui Qiu +6
E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are…
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Jia Liu, ChangYi He, YingQiao Lin +3
Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
Yiming Lei, Zhizheng Yang, Zeming Liu +5
Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal mode…
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
Yiming Lei, Chenkai Zhang, Zeming Liu +5
Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…