3 papers
cs.AI2026
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
Tu Trinh, Mohamed Elfeki, Guangze Luo +9
Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is not raw capability, but judgm…
cs.AI2026
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie +14
Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While ex…
cs.CV2025
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
Xingang Guo, Utkarsh Tyagi, Advait Gosai +8
Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such…