collaborators

9 papers

cs.LG2026

Qwen-CUA: Native Computer Use for (almost) Everything

Dunjie Lu, Shuai Bai, Tianyi Bai +42

Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive expe…

cs.SE2026

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation

Zhengran Zeng, Ruikai Shi, Keke Han +7

Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Languag…

cs.CV2025

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

Bofei Zhang, Zirui Shang, Zhi Gao +7

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. Ho…

cs.AI2025

MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems

Haifeng Jing, Yujie Hou, Junfei Liu +4

With the rapid development of Large Language Models, dialogue systems are shifting from information tools to emotional companions, heralding the era of Emotional Companionship Dial…

cs.CL2025

Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation

Bo Li, Zhenghua Xu, Rui Xie

Multilingual Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to perform knowledge-intensive tasks in multilingual settings by leveraging retrieved documen…

cs.SE2025

Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development

Zhengran Zeng, Yixin Li, Rui Xie +2

The development of LLM-based autonomous agents for end-to-end software development represents a significant paradigm shift in software engineering. However, the scientific evaluati…