1 citations · 1 across the 10 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Data to Behavior: Predicting Unintended Model Behaviors Before Training
Mengru Wang, Zhenqian Xu, Junfeng Fang +4
Large Language Models (LLMs) can acquire unintended biases from seemingly benign training data even without explicit cues or malicious content. Existing methods struggle to detect…
cs.LG2025
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
Yixin Ou, Yunzhi Yao, Ningyu Zhang +5
Despite exceptional capabilities in knowledge-intensive tasks, Large Language Models (LLMs) face a critical gap in understanding how they internalize new knowledge, particularly ho…