#interpretability

try —

24 papers match

cs.LG2026

Information Bottleneck Learning for Faithful Time Series Forecasting Explanations

Xu Zheng, Wei Cheng, Zhuomin Chen +3

The paper presents IB-Forecast, an interpretable multivariate time-series forecasting model that uses an information bottleneck to generate sparse, faithful explanations of predict…

#time series forecasting#interpretability#information bottleneck#explainable AI
cs.LG2026

TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees

James Chapman

TreeCCA introduces a method to train gradient-boosted tree ensembles as end-to-end CCA encoders using an Eckart‑Young loss, achieving nonlinear correlation extraction with native i…

#canonical correlation analysis#gradient-boosted trees#multiview learning#interpretability
cs.LG2026

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

Lei Dong

The paper proposes a driven‑nucleation rate law describing how language model capabilities emerge only when an entire circuit aligns in a single stochastic event, and shows that re…

#emergent capabilities#transformer circuit analysis#training dynamics#interpretability
cs.LG2026

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

Yixuan Duan, Wei Qiu

The paper introduces ECG-InterpBench, a benchmark that uses matched-capacity sparse autoencoders to assess how interpretable the internal representations of frozen ECG foundation m…

#ecg#foundation models#interpretability#sparse autoencoders
cs.AI2026

Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

Hyunkyung Han, Min Jung Kim

The paper investigates how loss invariance affects the encoding of volume concepts in a transformer-based echocardiography model, showing that without explicit volume supervision t…

#concept bottleneck models#echocardiography#volume estimation#interpretability
cond-mat.soft2026

Physics-Guided Interpretable Machine Learning Framework for Anomalous Transport in Crowded Media with Tunable Flexibility

Zakiya Shireen, Sujin B. Babu

The paper introduces a physics‑guided interpretable machine‑learning framework that combines Brownian Cluster Dynamics simulations with surrogate models and SHAP analysis to quanti…

#anomalous transport#crowded media#machine learning#interpretability
cs.LG2026

Counterfactuals for Feature-Weighted Clustering

Richard J. Fawley, Renato Cordeiro de Amorim

The paper proposes VoICE, a framework that generates counterfactual explanations for feature‑weighted k‑means clustering by projecting points onto weighted Voronoi regions of targe…

#counterfactual explanations#clustering#feature-weighted k-means#voronoi regions
cs.LG2026

TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation

Wen Yang Tan, Jiawei Li, Fang Liu +4

The paper introduces TIDE, a machine‑learning system that combines battery domain knowledge with operational data to estimate battery health accurately while providing trustworthy…

#battery health estimation#trustworthy AI#interpretability#symbolic distillation
cs.CR2026

The Refusal Residue: When Probes Catch Alignment Faking and When They Don't

Aman Mehta

The paper investigates whether hidden states of large language models can reveal when the model is faking compliance (alignment faking) and finds that detection is possible for som…

#alignment faking#model probing#refusal detection#large language models
econ.EM2026

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

Jacob Carlson

The paper proposes a framework that converts unstructured data into sparse, interpretable concept embeddings and then applies high‑dimensional multiple hypothesis testing with sele…

#high-dimensional inference#multiple hypothesis testing#interpretability#concept embeddings
cs.CV2026

Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification

Honglei Jin, Wenshuo Chen, Shaofeng Liang +6

The paper presents Traffic-CBM, a multimodal framework that classifies encrypted network traffic by converting flow statistics, temporal features, and byte-level data into hierarch…

#encrypted traffic classification#multimodal learning#interpretability#concept learning
cs.AI2026

From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

Ingmar Posner, Anson Lei, Bernhard Schölkopf

The paper introduces Mechanistic World Models, a framework that places reusable explanatory mechanisms at the core of AI systems to enable autonomous scientific discovery beyond me…

#mechanistic world models#scientific discovery#causal representation learning#modular architectures
cs.AI2026

AIMO Interpretability Challenge

Michal Štefánik, Philipp Mondorf, Andreas Waldis +11

The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…

#interpretability#robustness#mathematical reasoning#benchmarking
quant-ph2026

Inherent interpretability provides inherent value in quantum machine learning

Kaitlin Gili, Zachary P. Bradshaw

The paper argues that the intrinsic mathematical structure of quantum machine learning models can provide inherent interpretability, offering value beyond raw performance, and illu…

#interpretability#quantum Fourier models#gaussian processes#kernel design
cs.LG2026

Understanding Structured Health Data through Interaction-Aware Mixture-of-Experts

Ji Hwan Park, Ying Ding, Tianjin Guo

The paper proposes an interaction-aware mixture-of-experts model to predict post-stroke rigidity using multi-level views of structured health records, emphasizing interpretability…

#health data#mixture of experts#post-stroke prediction#interpretability
cs.CV2026

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

Jinjian Wu, Jiaqi Tang, Wei Wei +5

The paper introduces IQA-T1, a framework that combines multimodal large language models with specialized visual analysis tools to generate explicit evidence (e.g., noise residual m…

#image quality assessment#visual evidence reasoning#multimodal language models#tool-based analysis
cs.LG2026

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini

The paper proposes two sparsity regularizers that work with Top‑k sparse autoencoders to improve the interpretability of latent features without hurting reconstruction quality.

#sparse autoencoders#top‑k selection#regularization#interpretability
cs.RO2026

Unveiling Complex Collective Behaviors from Simple Rewards

Yize Mi, Jianan Li, Liang Li +1

The paper introduces an explanatory framework with an Agent Response Map to interpret how simple reward signals lead to complex collective behaviors in robot swarms, revealing hidd…

#multi-agent reinforcement learning#swarm robotics#interpretability#geometric field analysis
physics.ao-ph2026

Robustness of Deep Learning Models for PV Power Forecasting under NWP Forecast Errors: A Spatiotemporal and Physically Interpretable Analysis

Dandan Chen, Yan Zhao, Xuepeng Chen

The paper evaluates how deep learning and machine‑learning models for photovoltaic power forecasting behave when faced with realistic, temporally correlated errors in numerical wea…

#photovoltaic forecasting#model robustness#numerical weather prediction errors#deep sequence models
cs.SD2026

Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning

Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov +5

The paper introduces HIR‑SDD, a speech deepfake detection framework that leverages large audio language models and chain‑of‑thought reasoning from a human‑annotated dataset to impr…

#speech deepfake detection#robustness#interpretability#large audio language models
math.LO2026

Weak essentially undecidable theories of hereditarily finite multisets

Platon Sifnaios

The paper defines two first‑order theories for hereditarily finite multisets, shows they are mutually interpretable with Robinson's theories R and Q, and proves that the finitely a…

#first-order theory#hereditarily finite multisets#essential undecidability#interpretability
eess.SP2026

From Wireless SNNs to SN P Systems: A Low-Energy Rule-Based Conversion

Pietro Savazzi, Mauro Marchese, Anna Vizziello +1

The paper presents a method to transform trained distributed wireless spiking neural networks into rule‑based Spiking Neural P (SN P) systems, yielding interpretable, human‑readabl…

#spiking neural networks#sn p systems#energy‑efficient edge inference#interpretability
cs.LG2026

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Zixiang Xu, Sixian Li, Huaxing Liu +4

The paper investigates how biases in large language models used as judges are reflected in their hidden activations, identifying low-dimensional subspaces that encode bias and show…

#bias detection#interpretability#large language models#fairness
cs.LG2026

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

Ayush Karmacharya, Luke Luschwitz, Lucia Romero +2

The paper proposes using sparse autoencoders to extract interpretable sparse features from intermediate neural network layers and defines an OOD detection score based on cosine sim…

#out-of-distribution detection#sparse autoencoders#interpretability#representation learning

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.