activity
20242026
collaborators

7 papers

cs.AI2026

When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

Samuel Jacob Chacko, James Hugglestone, Chashi Mahiul Islam +1

Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by an average of 16.2~percentag…

cs.AI2026

Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models

Chashi Mahiul Islam, Alan Villarreal, Mao Nishino +2

As Large Language Models (LLMs) are increasingly integrated into agentic workflows, their unpredictability stemming from numerical instability has emerged as a critical reliability…

cs.CV2025

Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning

Chashi Mahiul Islam, Oteo Mamo, Samuel Jacob Chacko +2

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we…

cs.CV2025

Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging

Montasir Shams, Chashi Mahiul Islam, Shaeke Salman +2

Vision transformers (ViTs) have rapidly gained prominence in medical imaging tasks such as disease classification, segmentation, and detection due to their superior accuracy compar…

cs.CV2025

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

Chashi Mahiul Islam, Samuel Jacob Chacko, Preston Horne +1

Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alternative offering competitive perfor…

cs.CV2025

Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers

Chashi Mahiul Islam, Samuel Jacob Chacko, Mao Nishino +1

While transformer-based models dominate NLP and vision applications, their underlying mechanisms to map the input space to the label space semantically are not well understood. In…