activity
20242026
collaborators

6 papers

cs.AI2026

WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark

Peng Yuan, Yuyang Yin, Yuxuan Cai +1

Existing browser agent benchmarks face a fundamental trilemma: real-website benchmarks lack reproducibility due to content drift, controlled environments sacrifice realism by omitt…

cs.CL2026

Yunque DeepResearch Technical Report

Yuxuan Cai, Xinyi Lai, Peng Yuan +8

Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full…

cs.LG2025

FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction

Yuxuan Cai, Xiaozhuan Liang, Xinghua Wang +7

As large language models (LLMs) become increasingly powerful, the sequential nature of autoregressive generation creates a fundamental throughput bottleneck that limits the practic…

cs.CV2025

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning

Junzhe Xu, Yuyang Yin, Xi Chen

This paper introduces TBAC-UniImage, a novel unified model for multimodal understanding and generation. We achieve this by deeply integrating a pre-trained Diffusion Model, acting…

cs.CL2025

ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

Kangwei Liu, Siyuan Cheng, Bozhong Tian +7

Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the ov…

cs.CV2024

Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking

Zhengfei Xu, Sijia Zhao, Yanchao Hao +6

Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Pre…