works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.LG2026

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

Yanyu Ren, Xizheng Wang, Xiao Liu +8

Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with gro…

cs.AI2026

An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding

Yanyu Ren, Yunfeng Bai, Xizheng Wang +2

The paper presents MSEval, a benchmark that evaluates how multi‑agent coding systems build real‑world software, measuring functional success, latency, and token cost while varying…

cs.AI2026

EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents

Ying Mo, Yu Bai, Dapeng Sun +4

Recent advances in Multimodal Large Language Models (MLLMs) have enabled agents to operate in open-ended web and operating system environments. However, existing benchmarks predomi…

cs.AI2025

DMA: Online RAG Alignment with Human Feedback

Yu Bai, Yukai Miao, Dawei Wang +9

Retrieval-augmented generation (RAG) systems often rely on static retrieval, limiting adaptation to evolving intent and content drift. We introduce Dynamic Memory Alignment (DMA),…

cs.SE2025

The CodeInverter Suite: Structure- and Data-Aware Binary Decompilation with Efficient LLMs

Peipei Liu, Jian Sun, Rongkang Sun +7

Binary decompilation plays a vital role in various cybersecurity and software engineering tasks. Recently, end-to-end decompilation methods powered by large language models (LLMs)…

cs.CR2025

The Digital Cybersecurity Expert: How Far Have We Come?

Dawei Wang, Geng Zhou, Xianglong Li +5

The increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evalua…