7 papers
Concept-SAE: A Controllable and Invertible Concept Interface for Sparse Autoencoders
Jianrong Ding, Muxi Chen, Chenchen Zhao +1
Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, providing a powerful lens for passive feature discovery. However, this passive…
GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval
Zhaohua Zhang, Jianhuan Zhuo, Muxi Chen +8
The CLIP model has established itself as a cornerstone of large-scale retrieval systems. However, its performance often degrades under distributional shifts such as multilingual, l…
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
Chenchen Zhao, Jianhuan Zhuo, Muxi Chen +6
Composed image retrieval (CIR) requires multi-modal models to jointly reason over visual content and semantic modifications presented in text-image input pairs. While current CIR m…
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
Chenchen Zhao, Muxi Chen, Qiang Xu
Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on w…
FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
Muxi Chen, Zhaohua Zhang, Chenchen Zhao +8
Static benchmarks have provided a valuable foundation for comparing Text-to-Image (T2I) models. However, their passive design offers limited diagnostic power, struggling to uncover…
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen +19
The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehens…