8 papers
Advancing Creative Physical Intelligence in Large Multimodal Models
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded…
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains under…
OSExpert: Computer-Use Agents Learning Professional Skills via Exploration
Jiateng Liu, Zhenhailong Wang, Rushi Wang +6
General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less…
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
Sethuraman T, Savya Khosla, Aditi Tiwari +11
This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…
Predicting Camera Pose from Perspective Descriptions for Spatial Reasoning
Xuejun Zhang, Aditi Tiwari, Zhenhailong Wang +1
Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views…
Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos
Aditi Tiwari, Farzaneh Masoud, Dac Trong Nguyen +3
Modern AI systems struggle most in environments where reliability is critical - scenes with smoke, poor visibility, and structural deformation. Each year, tens of thousands of fire…