works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

cs.LG2026

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

Sriram Balasubramanian, Soheil Feizi

Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training…

cs.CV2026

SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs

Parsa Hosseini, Sumit Nawathe, Mazda Moayeri +2

The paper introduces SpurLens, an automated pipeline that uses GPT-4 and open-set object detectors to find spurious visual cues in multimodal large language models, showing that th…

cs.CL2025

A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models

Sriram Balasubramanian, Samyadeep Basu, Soheil Feizi

Chain-of-thought (CoT) reasoning enhances performance of large language models, but questions remain about whether these reasoning traces faithfully reflect the internal processes…

cs.AI2025

Tool Preferences in Agentic LLMs are Unreliable

Kazem Faghih, Wenxiao Wang, Yize Cheng +5

Large language models (LLMs) can now access a wide range of external tools, thanks to the Model Context Protocol (MCP). This greatly expands their abilities as various agents. Howe…

cs.CL2025

Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis

Anushka Yadav, Isha Nalawade, Srujana Pillarichety +7

The emergence of reasoning models and their integration into practical AI chat bots has led to breakthroughs in solving advanced math, deep search, and extractive question answerin…

cs.LG2025

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Zihao Lin, Samyadeep Basu, Mohammad Beigi +18

The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…