activity
20242026
collaborators

11 papers

cs.CV2026

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

Zichen Zhu, Yuheng Sun, Mingxuan Zhu +11

Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes. Creations by generative models may contain…

cs.AI2026

No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

Zixu Yang, Hang Zheng, Nan Jiang +5

Large language model (LLM) agents have increasingly advanced service applications, such as booking flight tickets. However, these service agents suffer from unreliability in long-h…

cs.CE2026

ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge

Zihan Zhao, Ziping Wan, Lu Chen +12

Atomized chemical knowledge, such as functional group information of molecules and reactions, plays a pivotal intermediate role in the reasoning process that connects molecular str…

cs.AI2026

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

Situo Zhang, Yifan Zhang, Zichen Zhu +6

Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging for multimodal large language models (MLLM…

cs.CL2026

PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length

Situo Zhang, Yifan Zhang, Zichen Zhu +5

Speculative decoding (SD) is a powerful technique for accelerating the inference process of large language models (LLMs) without sacrificing accuracy. Typically, SD employs a small…

cs.CL2025

MULTI: Multimodal Understanding Leaderboard with Text and Images

Zichen Zhu, Yang Xu, Lu Chen +11

The rapid development of multimodal large language models (MLLMs) raises the question of how they compare to human performance. While existing datasets often feature synthetic or o…