activity
20182026
most citedEnhancing Exchange Rate Forecasting with Explainable Deep Learning Models

2 citations · 2 across the 16 of their papers we have counts for

collaborators

26 papers

cs.AI2026

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

Chenqian Le, Jiayi Cheng, Qijia He +3

Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and co…

cs.CL2026

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

Sky Ng, Brihi Joshi, Ishan Gupta +47

Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain i…

cs.HC2026

PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

Yifan Simon Liu, Qianfeng Wen, Yilan Fan +40

Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and…

cs.AI2026

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Qijia He, Jiayi Cheng, Chenqian Le +8

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware system…

cs.AI2026

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

Zhengtao Yao, Runhao Li, Xupeng Chen +12

Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretr…

cs.AI2026

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Zhengtao Yao, Runhao Li, Xupeng Chen +12

Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can…