collaborators

12 papers

cs.CV2026

DanceOPD: On-Policy Generative Field Distillation

Wei Zhou, Xiongwei Zhu, Zelin Xu +8

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are…

cs.CV2026

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs

Chengwen Liu, Hao Peng, Jisheng Dang +3

In multimodal video reasoning, reinforcement learning-based methods typically rely on simplistic and inflexible reasoning-length control strategies that fail to adapt to the model'…

cs.CL2026

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play

Leyang Shen, Yang Zhang, Xiaoyan Zhao +2

Large language model (LLM)-based multi-agent systems (MAS) have demonstrated great potential in solving tasks with execution complexity, by distributing subtasks across cooperative…

cs.CV2026

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs

Chengwen Liu, Zhe Huang, Jisheng Dang +3

Reinforcement learning has improved the reasoning ability of large language models, but applying outcome-only rewards to video multimodal large language models (Video-MLLMs) provid…

cs.LG2026

CARL: Criticality-Aware Agentic Reinforcement Learning

Leyang Shen, Yang Zhang, Chun Kai Ling +2

Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction. However, in such multi-step settings,…

cs.CV2026

DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

Jiawei Wang, Ming Lei, Yaning Yang +8

Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biod…