collaborators

5 papers

cs.LG2026

Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards

Yingyu Shan, Yuhang Guo, Zihao Cheng +7

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform c…

cs.CL2026

PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes

Yingyu Shan, Zeming Liu, Silin Li +4

Recent advancements in Large Language Models (LLMs) have empowered home assistants with natural language interaction capabilities. However, current assistants overlook the progress…

cs.CL2026

TrustMargin: Training-Free Arbitration between Parametric Memory and Retrieved Evidence in Large Language Models

Jingyan Xu, Hong Shi, Yi Shan +4

Large language models answer knowledge-intensive questions using both parametric memory and retrieved evidence, but neither source is uniformly reliable. Retrieval can fill knowled…

cs.CL2026

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

XiuYu Zhang, Yi Shan, Junfeng Fang +1

Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is…

cs.CL2024

FAME: Towards Factual Multi-Task Model Editing

Li Zeng, Yingyu Shan, Zeming Liu +2

Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks. Nevertheless, outdated knowledge or factual errors within…