collaborators

5 papers

cs.AI2026

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges

Zi-Han Wang, Lam Nguyen, Zhengyang Zhao +4

The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success o…

cs.CL2026

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Xinyi Li, Zhen Fang, Yongxin Deng +12

Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challenges: inconsistent inference con…

cs.LG2026

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

Ruiying Peng, Mengyu Yang, Jing Lei +3

Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is no…

cs.LG2026

Causal Fine-Tuning under Latent Confounded Shift

Jialin Yu, Yuxiang Zhou, Haoxuan Li +6

Adapting to latent confounded shift remains a core challenge in modern AI. This setting is driven by hidden variables that induce spurious correlations between inputs and outputs d…

cs.LG2024

When Can Proxies Improve the Sample Complexity of Preference Learning?

Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4

We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…