Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
Haowen Gao, Zhenyu Zhang, Liang Pang +7
Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…
cs.AI2025
Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
Chenyu Zhou, Xiaoming Shi, Hui Qiu +6
E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are…