15 citations · 24 across the 12 of their papers we have counts for
12 papers
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
Tianyu Guo, Druv Pai, Yu Bai +3
Practitioners have consistently observed three puzzling phenomena in transformer-based large language models (LLMs): attention sinks, value-state drains, and residual-state peaks,…
CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models
Ying Nie, Binwei Yan, Tianyu Guo +9
Large language models (LLMs) have achieved remarkable performance on various NLP tasks, yet their potential in more challenging and domain-specific task, such as finance, has not b…
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
Danni Yang, Jiayi Ji, Yiwei Ma +4
In this paper, we introduce SemiRES, a semi-supervised framework that effectively leverages a combination of labeled and unlabeled data to perform RES. A significant hurdle in appl…
Collaborative Heterogeneous Causal Inference Beyond Meta-analysis
Tianyu Guo, Sai Praneeth Karimireddy, Michael I. Jordan
Collaboration between different data centers is often challenged by heterogeneity across sites. To account for the heterogeneity, the state-of-the-art method is to re-weight the co…
A robust audio deepfake detection system via multi-view feature
Yujie Yang, Haochen Qin, Hang Zhou +4
With the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio de…
Data-Free Distillation of Language Model by Text-to-Text Transfer
Zheyuan Bai, Xinduo Liu, Hailin Hu +3
Data-Free Knowledge Distillation (DFKD) plays a vital role in compressing the model when original training data is unavailable. Previous works for DFKD in NLP mainly focus on disti…