From the 2 of 29 linked papers with an AI index.
29 papers
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
Bowei He, Yankai Chen, Xiaokun Zhang +1
The paper proposes Branching Policy Optimization (BPO), a reinforcement learning method for large language model agents operating in deterministic, snapshottable sandboxes, which l…
Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
Ye Yuan, Weien Li, Rui Song +19
The paper proposes a unified framework for discrete denoising diffusion models that ties together tokenization, vocabulary design, and generation methods, showing how existing appr…
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
Ãdám Kovács, Bowei He, Xue Liu +3
Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generation systems increasingly rely…
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
Yankai Chen, Hanrong Zhang, Bowei He +2
Standard Set Representation Learning methods typically excel on curated data but often overlook the challenge of inference-time element corruption. This refers to scenarios where d…
QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
Ye Yuan, Rui Song, Weien Li +12
Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environ…
Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization
Yonghan Yang, Ye Yuan, Zipeng Sun +5
Offline black-box optimization aims to discover novel designs with high property scores using only a static dataset, a task fundamentally challenged by the out-of-distribution (OOD…