15 papers
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
Lu Jia, Haibo Tong, Feifei Zhao +5
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments,…
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
Guobin Shen, Lei Huang, Xiang Cheng +4
On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback to serve as its own teacher…
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
Guobin Shen, Xiang Cheng, Chenxiao Zhao +4
On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising directi…
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
Jinchang Zhu, Jindong Li, Chengyu Zou +4
Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effecti…
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Jinchang Zhu, Jindong Li, Yuwen Hao +3
A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraining: upper layers commit to s…
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
Haibo Tong, Feifei Zhao, Linghao Feng +18
Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control…