works on

From the 1 of 35 linked papers with an AI index.

activity
20242026
most citedReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Multimodal Concept Bottleneck Models

Tongqing Shi, Ge Yan, Tuomas Oikarinen +1

Concept Bottleneck Models (CBMs) enhance the interpretability of deep learning networks by aligning the features extracted from images with natural concepts. However, existing CBMs…

cs.CV2025

Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability

Tuomas Oikarinen, Ge Yan, Akshay Kulkarni +1

Interpreting individual neurons or directions in activation space is an important topic in mechanistic interpretability. Numerous automated interpretability methods have been propo…

cs.CV2025

Interpretable Generative Models through Post-hoc Concept Bottlenecks

Akshay Kulkarni, Ge Yan, Chung-En Sun +2

Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to des…

cs.CV2025

RAT: Boosting Misclassification Detection Ability without Extra Data

Ge Yan, Tsui-Wei Weng

As deep neural networks(DNN) become increasingly prevalent, particularly in high-stakes areas such as autonomous driving and healthcare, the ability to detect incorrect predictions…

cs.CV2025

Interpreting Neurons in Deep Vision Networks with Language Models

Nicholas Bai, Rahul A. Iyer, Tuomas Oikarinen +2

In this paper, we propose Describe-and-Dissect (DnD), a novel method to describe the roles of hidden neurons in vision networks. DnD utilizes recent advancements in multimodal deep…

cs.CV2025

VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance

Divyansh Srivastava, Ge Yan, Tsui-Wei Weng

Concept Bottleneck Models (CBMs) provide interpretable prediction by introducing an intermediate Concept Bottleneck Layer (CBL), which encodes human-understandable concepts to expl…