Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks
Tianze Han, Beining Xu, Hanbo Zhang +1
Mental-health dialogue models are increasingly evaluated by AI-based evaluators, yet these evaluators often treat surface empathy, supportiveness, or fluency as evidence of safety.…
cs.CL2025
TECP: Token-Entropy Conformal Prediction for LLMs
Beining Xu, Yongming Lu
Uncertainty quantification (UQ) for open-ended language generation remains a critical yet underexplored challenge, especially under black-box constraints where internal model signa…
cs.CL2025
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
Beining Xu, Arkaitz Zubiaga
Large Language Models (LLMs) have demonstrated exceptional performance on a range of downstream NLP tasks by generating text that closely resembles human writing. However, the ease…