5 papers
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
Qi Sun, Siyue Zhang, Yulin Chen +3
Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data from weak-weaker model pairs (e…
Causal Discovery as Dialectical Aggregation: A Quantitative Argumentation Framework
Sheng Wei, Yulin Chen, Beishui Liao
Constraint-based causal discovery is brittle in finite-sample regimes because erroneous conditional-independence (CI) decisions can cascade into substantial structural errors. We p…
Search, Do not Guess: Teaching Small Language Models to Be Effective Search Agents
Yizhou Liu, Qi Sun, Yulin Chen +2
Agents equipped with search tools have emerged as effective solutions for knowledge-intensive tasks. While Large Language Models (LLMs) exhibit strong reasoning capabilities, their…
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
Chen Yueh-Han, Nitish Joshi, Yulin Chen +3
Current LLM safety defenses fail under decomposition attacks, where a malicious goal is decomposed into benign subtasks that circumvent refusals. The challenge lies in the existing…
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
Anqi Zhang, Yulin Chen, Jane Pan +4
Reasoning models have achieved remarkable performance on tasks like math and logical reasoning thanks to their ability to search during reasoning. However, they still suffer from o…