7 papers
Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records
Fabian Degen, Oishi Deb, Jindong Gu +4
Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We intr…
Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa +8
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
Oishi Deb, Anjun Hu, Ashkan Khakzar +2
We propose a training-free method, Articulate3D, to pose a 3D asset through language control. Despite advances in vision and language models, this task remains surprisingly challen…
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
Qianyi Deng, Oishi Deb, Amir Patel +4
Animal pose estimation (APE) aims to locate the animal body parts using a diverse array of sensor and modality inputs (e.g. RGB cameras, LiDAR, infrared, IMU, acoustic and language…
Responsible AI Governance: A Response to UN Interim Report on Governing AI for Humanity
Sarah Kiden, Bernd Stahl, Beverley Townsend +22
This report presents a comprehensive response to the United Nation's Interim Report on Governing Artificial Intelligence (AI) for Humanity. It emphasizes the transformative potenti…