collaborators

8 papers

cs.AI2026

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

Haowen Gao, Zhenyu Zhang, Liang Pang +7

Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…

cs.CV2025

Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering

Jun Li, Hongjian Dou, Zhenyu Zhang +3

Composed Image Retrieval (CIR) presents a significant challenge as it requires jointly understanding a reference image and a modified textual instruction to find relevant target im…

cs.AI2025

Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules

Chenyu Zhou, Xiaoming Shi, Hui Qiu +6

E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are…

cs.LG2025

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Jia Liu, ChangYi He, YingQiao Lin +3

Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…

cs.CL2025

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Yiming Lei, Zhizheng Yang, Zeming Liu +5

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal mode…

cs.CL2025

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

Yiming Lei, Chenkai Zhang, Zeming Liu +5

Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…