16 papers · 1 filter
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
Md Abrar Jahin, Md Rizwan Parvez
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language…
A Survey on Agentic Security: Applications, Threats and Defenses
Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed +2
LLM-based agents are now used throughout cybersecurity. While these agents facilitate powerful and autonomous security applications, their autonomy opens up new attack surfaces, an…
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
Raman Saparkhan, Majd Hawasly, Md Rizwan Parvez +1
Self-consistency (SC) is a popular technique for improving the reasoning accuracy of large language models by aggregating multiple sampled outputs, but it comes at a high computati…
AI Debate Aids Assessment of Controversial Claims
Salman Rahman, Sheriff Issaka, Ashima Suvarna +11
As AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides-…
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam +8
Dashboards are powerful visualization tools for data-driven decision-making, integrating multiple interactive views that allow users to explore, filter, and navigate data. Unlike s…
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team
Md Tanzib Hosain, Salman Rahman, Md Kishor Morol +1
Despite impressive progress on complex reasoning, current large language models (LLMs) typically operate in isolation - treating each problem as an independent attempt, without acc…