5 papers
PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
Syed Nazmus Sakib, Nafiul Haque, Mohammad Zabed Hossain +1
Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis. To address this, we pr…
PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
Nafiul Haque, Syed Nazmus Sakib, Shifat E Arman
Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, they systematically hallucinat…
The Surface You Test Is Not the Surface That Breaks
Shifat E Arman, Syed Nazmus Sakib, Nafiul Haque +1
Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the agent then executes as if they…
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin +4
Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-based adaptive questioning. In pl…
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
Shifat E. Arman, Syed Nazmus Sakib, Tapodhir Karmakar Taton +2
We introduce PATHWAYS, a benchmark of 250 multi-step decision tasks that test whether web-based agents can discover and correctly use hidden contextual information. Across both clo…