84 citations · 139 across the 11 of their papers we have counts for
Showing 2025Show all
3 papers · 1 filter
q-bio.QM2025
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
Xiangyu Liu, Haodi Lei, Yi Liu +2
Sparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming…
cs.SE2025
Investigating Training Data Detection in AI Coders
Tianlin Li, Yunxiang Wei, Zhiming Li +5
Recent advances in code large language models (CodeLLMs) have made them indispensable tools in modern software engineering. However, these models occasionally produce outputs that…
cs.CV2025
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
Xiaojun Bi, Shuo Li, Junyao Xing +7
Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to th…