collaborators

8 papers

cs.AI2026

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

Haowen Gao, Zhenyu Zhang, Liang Pang +7

Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…

cs.AI2025

Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules

Chenyu Zhou, Xiaoming Shi, Hui Qiu +6

E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are…

cs.LG2025

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Jia Liu, ChangYi He, YingQiao Lin +3

Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…

cs.CV2025

Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering

Jun Li, Hongjian Dou, Zhenyu Zhang +3

Composed Image Retrieval (CIR) presents a significant challenge as it requires jointly understanding a reference image and a modified textual instruction to find relevant target im…

cs.CL2025

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

Yiming Lei, Chenkai Zhang, Zeming Liu +5

Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…

cs.CV2025

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

Chenkai Zhang, Yiming Lei, Zeming Liu +5

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities o…