2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.AI2026★ 2 cited
General Agent Evaluation
Elron Bandel, Asaf Yehudai, Lilach Eden +12
General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes…
cs.CL2026
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
Tomer Ashuach, Shai Gretz, Yoav Katz +2
Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether large language models possess si…