activity
20242026
most citedMulti-Programming Language Sandbox for LLMs

1 citations · 1 across the 11 of their papers we have counts for

collaborators

20 papers

cs.CV2026

State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

Mingxu Chai, Chenyu Liu, Ziyu Shen +7

Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approach…

cs.LG2026

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

Bing Shao, Jiazheng Zhang, Long Ma +11

On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens rem…

cs.AI2026

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Boyang Liu, Senjie Jin, Peixin Wang +15

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those…

cs.CL2026

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

Dingwei Zhu, Jiahan Li, Chengjun Pan +22

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history sca…

cs.LG2026

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

Jiawei Zheng, Jiazhen Zhang

Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assum…

cs.AI2026

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

Zhiheng Xi, Dingwen Yang, Jiaqi Liu +21

Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic eval…