5 papers
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
Jun Rui Huang, Wang Bill Zhu, Ziyi Liu +3
Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these in…
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
Wang Bill Zhu, Miaosen Chai, Shangshang Wang +5
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during de…
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
Wang Bill Zhu, Qiutong Tony Yi, Robin Jia +1
Large language models (LLMs) perform substantially below human level on existing theory-of-mind (ToM) benchmarks, even when augmented with chain-of-thought prompting or probabilist…
Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks
Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3
As LLM-based assistants become persistent and personalized, they must extract and retain useful information from past conversations as memory. However, the types of information wor…
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
Wang Bill Zhu, Tianqi Chen, Xinyan Velocity Yu +6
Cancer patients are increasingly turning to large language models (LLMs) for medical information, making it critical to assess how well these models handle complex, personalized qu…