2 papers
cs.LG2026
Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models
Aryan Khurana, Aravind Ramana RN, Dhruv Kumar
Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly…
cs.LG2025
Benchmarking the Generality of Vision-Language-Action Models
Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka +6
Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices r…