6 papers
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
Yearim Kim, Sangyu Han, Nojun Kwak
Modern vision models achieve remarkable accuracy, but explaining where evidence arises, what the model encodes, and how internal computations assemble that evidence remains fragmen…
VDPP: Video Depth Post-Processing for Speed and Scalability
Daewon Yoon, Injun Baek, Sangyu Han +2
Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end-to-end video depth models have es…
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
Injun Baek, Yearim Kim, Nojun Kwak
While advancements in Text-to-Video (T2V) generative AI offer a promising path toward democratizing content creation, current models are often optimized for visual fidelity rather…
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
Jinyung Hong, Yearim Kim, Keun Hee Park +3
Inner interpretability is a promising field aiming to uncover the internal mechanisms of AI systems through scalable, automated methods. While significant research has been conduct…
Causal Interpretation of Sparse Autoencoder Features in Vision
Sangyu Han, Yearim Kim, Nojun Kwak
Understanding what sparse auto-encoder (SAE) features in vision transformers truly represent is usually done by inspecting the patches where a feature's activation is highest. Howe…
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
Sangyu Han, Yearim Kim, Nojun Kwak
The truthfulness of existing explanation methods in authentically elucidating the underlying model's decision-making process has been questioned. Existing methods have deviated fro…