2 papers
cs.CL2025
Consistency of Responses and Continuations Generated by Large Language Models on Social Media
Wentao Xu, Wenlu Fan, Yuqi Zhu +1
Large Language Models (LLMs) demonstrate remarkable capabilities in text generation, yet their emotional consistency and semantic coherence in social media contexts remain insuffic…
cs.CL2025
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…