#interpretability
24 papers match
Information Bottleneck Learning for Faithful Time Series Forecasting Explanations
Xu Zheng, Wei Cheng, Zhuomin Chen +3
The paper presents IB-Forecast, an interpretable multivariate time-series forecasting model that uses an information bottleneck to generate sparse, faithful explanations of predict…
TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees
James Chapman
TreeCCA introduces a method to train gradient-boosted tree ensembles as end-to-end CCA encoders using an Eckart‑Young loss, achieving nonlinear correlation extraction with native i…
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
Lei Dong
The paper proposes a driven‑nucleation rate law describing how language model capabilities emerge only when an entire circuit aligns in a single stochastic event, and shows that re…
ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders
Yixuan Duan, Wei Qiu
The paper introduces ECG-InterpBench, a benchmark that uses matched-capacity sparse autoencoders to assess how interpretable the internal representations of frozen ECG foundation m…
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Hyunkyung Han, Min Jung Kim
The paper investigates how loss invariance affects the encoding of volume concepts in a transformer-based echocardiography model, showing that without explicit volume supervision t…
Physics-Guided Interpretable Machine Learning Framework for Anomalous Transport in Crowded Media with Tunable Flexibility
Zakiya Shireen, Sujin B. Babu
The paper introduces a physics‑guided interpretable machine‑learning framework that combines Brownian Cluster Dynamics simulations with surrogate models and SHAP analysis to quanti…
Counterfactuals for Feature-Weighted Clustering
Richard J. Fawley, Renato Cordeiro de Amorim
The paper proposes VoICE, a framework that generates counterfactual explanations for feature‑weighted k‑means clustering by projecting points onto weighted Voronoi regions of targe…
TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation
Wen Yang Tan, Jiawei Li, Fang Liu +4
The paper introduces TIDE, a machine‑learning system that combines battery domain knowledge with operational data to estimate battery health accurately while providing trustworthy…
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Aman Mehta
The paper investigates whether hidden states of large language models can reveal when the model is faking compliance (alignment faking) and finds that detection is possible for som…
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach
Jacob Carlson
The paper proposes a framework that converts unstructured data into sparse, interpretable concept embeddings and then applies high‑dimensional multiple hypothesis testing with sele…
Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification
Honglei Jin, Wenshuo Chen, Shaofeng Liang +6
The paper presents Traffic-CBM, a multimodal framework that classifies encrypted network traffic by converting flow statistics, temporal features, and byte-level data into hierarch…
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
Ingmar Posner, Anson Lei, Bernhard Schölkopf
The paper introduces Mechanistic World Models, a framework that places reusable explanatory mechanisms at the core of AI systems to enable autonomous scientific discovery beyond me…
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…
Inherent interpretability provides inherent value in quantum machine learning
Kaitlin Gili, Zachary P. Bradshaw
The paper argues that the intrinsic mathematical structure of quantum machine learning models can provide inherent interpretability, offering value beyond raw performance, and illu…
Understanding Structured Health Data through Interaction-Aware Mixture-of-Experts
Ji Hwan Park, Ying Ding, Tianjin Guo
The paper proposes an interaction-aware mixture-of-experts model to predict post-stroke rigidity using multi-level views of structured health records, emphasizing interpretability…
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Jinjian Wu, Jiaqi Tang, Wei Wei +5
The paper introduces IQA-T1, a framework that combines multimodal large language models with specialized visual analysis tools to generate explicit evidence (e.g., noise residual m…
Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders
Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini
The paper proposes two sparsity regularizers that work with Top‑k sparse autoencoders to improve the interpretability of latent features without hurting reconstruction quality.
Unveiling Complex Collective Behaviors from Simple Rewards
Yize Mi, Jianan Li, Liang Li +1
The paper introduces an explanatory framework with an Agent Response Map to interpret how simple reward signals lead to complex collective behaviors in robot swarms, revealing hidd…
Robustness of Deep Learning Models for PV Power Forecasting under NWP Forecast Errors: A Spatiotemporal and Physically Interpretable Analysis
Dandan Chen, Yan Zhao, Xuepeng Chen
The paper evaluates how deep learning and machine‑learning models for photovoltaic power forecasting behave when faced with realistic, temporally correlated errors in numerical wea…
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov +5
The paper introduces HIR‑SDD, a speech deepfake detection framework that leverages large audio language models and chain‑of‑thought reasoning from a human‑annotated dataset to impr…
Weak essentially undecidable theories of hereditarily finite multisets
Platon Sifnaios
The paper defines two first‑order theories for hereditarily finite multisets, shows they are mutually interpretable with Robinson's theories R and Q, and proves that the finitely a…
From Wireless SNNs to SN P Systems: A Low-Energy Rule-Based Conversion
Pietro Savazzi, Mauro Marchese, Anna Vizziello +1
The paper presents a method to transform trained distributed wireless spiking neural networks into rule‑based Spiking Neural P (SN P) systems, yielding interpretable, human‑readabl…
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Zixiang Xu, Sixian Li, Huaxing Liu +4
The paper investigates how biases in large language models used as judges are reflected in their hidden activations, identifying low-dimensional subspaces that encode bias and show…
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
Ayush Karmacharya, Luke Luschwitz, Lucia Romero +2
The paper proposes using sparse autoencoders to extract interpretable sparse features from intermediate neural network layers and defines an OOD detection score based on cosine sim…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.