3 papers
cs.AI2026
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang +7
Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likel…
cs.CL2025
Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models
Hanzhi Zhang, Sumera Anjum, Heng Fan +3
Hallucinations in generative AI, particularly in Large Language Models (LLMs), pose a significant challenge to the reliability of multilingual applications. Existing benchmarks for…
cs.SE2024
S3LLM: Large-Scale Scientific Software Understanding with LLMs using Source, Metadata, and Document
Kareem Shaik, Dali Wang, Weijian Zheng +4
The understanding of large-scale scientific software poses significant challenges due to its diverse codebase, extensive code length, and target computing architectures. The emerge…