10 papers
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Mingyue Cui, Linghui Shen, Xingyi Yang
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that…
Language-based Trial and Error Falls Behind in the Era of Experience
Haoyu Wang, Guozheng Ma, Shugang Cui +7
While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic or spatial tasks) remains limite…
Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
Yanming Li, Xuelin Zhang, WenJie Lu +11
Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems. However, training…
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Liang Chen, Xueting Han, Li Shen +2
Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the reasoning ability of large lang…
DeContext as Defense: Safe Image Editing in Diffusion Transformers
Linghui Shen, Mingyue Cui, Xingyi Yang
In-context diffusion models allow users to modify images with remarkable ease and realism. However, the same power raises serious privacy concerns: personal images can be easily ma…
Provably Robust Adaptation for Language-Empowered Foundation Models
Yuni Lai, Xiaoyu Xue, Linghui Shen +5
Language-empowered foundation models (LeFMs), such as CLIP and GraphCLIP, have transformed multimodal learning by aligning visual (or graph) features with textual representations,…