2 papers
cs.AI2026
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Tanmay Asthana, Aman Saksena, Divyansh Sahu
Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks measure factual recall, single-hop QA, or…
cs.CV2026
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects?
Animesh Maheshwari, Divyansh Sahu, Nishit Verma
Vision-language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting th…