Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Code Comprehension then Auditing for Unsupervised LLM Evaluation
Bhrij Patel, Souradip Chakraborty, Mengdi Wang +2
Large Language Models (LLMs) for unsupervised code correctness evaluation have recently gained attention because they can judge if code runs as intended without requiring reference…
cs.AI2025
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person…