Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
Haoran Wang, Maryam Khalid, Qiong Wu +2
Large language models (LLMs) are increasingly used in applications requiring factual accuracy, yet their outputs often contain hallucinated responses. While fact-checking can mitig…
cs.CL2025
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
Wendong Xu, Jing Xiong, Chenyang Zhao +16
We present SwingArena, a competitive evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static…