2 papers
cs.CL2025
Existing LLMs Are Not Self-Consistent For Simple Tasks
Zhenru Lin, Jiawen Tao, Yang Yuan +1
Large Language Models (LLMs) have grown increasingly powerful, yet ensuring their decisions remain transparent and trustworthy requires self-consistency -- no contradictions in the…
cs.AI2024
CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
Zhenru Lin, Yiqun Yao, Yang Yuan
Large language models (LLMs) such as ChatGPT are increasingly proficient in understanding and generating a mixture of code and text. Evaluation based on such can…