collaborators

5 papers

cs.AI2026

APeB: Benchmarking Personalization Ability of Large Language Model Agents

Garry Yang, Zizhe Chen, Xinru Chen +9

LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy inte…

cs.LG2026

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Zizhe Chen, Jiqian Dong, Yizhou Tian +4

Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state value estimation is critical for…

cs.LG2026

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

Zixuan Chen, Hao Lin, Zizhe Chen +6

LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply rather than correct. We term this…

cs.CL2026

OpenAI GPT-5 System Card

Aaditya Singh, Adam Fry, Adam Perelman +483

This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…

cs.CV2025

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

Garry Yang, Zizhe Chen, Man Hon Wong +5

Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic vid…