activity
20242026
collaborators

7 papers

cs.LG2026

Concept-SAE: A Controllable and Invertible Concept Interface for Sparse Autoencoders

Jianrong Ding, Muxi Chen, Chenchen Zhao +1

Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, providing a powerful lens for passive feature discovery. However, this passive…

cs.CV2026

FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval

Chenchen Zhao, Jianhuan Zhuo, Muxi Chen +6

Composed image retrieval (CIR) requires multi-modal models to jointly reason over visual content and semantic modifications presented in text-image input pairs. While current CIR m…

cs.CV2026

\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions

Chenchen Zhao, Muxi Chen, Qiang Xu

Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on w…

cs.CV2025

FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration

Muxi Chen, Zhaohua Zhang, Chenchen Zhao +8

Static benchmarks have provided a valuable foundation for comparing Text-to-Image (T2I) models. However, their passive design offers limited diagnostic power, struggling to uncover…

cs.CV2025

HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging

Muxi Chen, Chenchen Zhao, Qiang Xu

Despite the significant success of deep learning models in computer vision, they often exhibit systematic failures on specific data subsets, known as error slices. Identifying and…

cs.CL2025

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities

MediaTek Research, :, Chan-Jan Hsu +6

Llama-Breeze2 (hereinafter referred to as Breeze2) is a suite of advanced multi-modal language models, available in 3B and 8B parameter configurations, specifically designed to enh…