21 papers
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
Yinuo Jiang, Yongjie Ye, Zhou Tao +4
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve…
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Yuqi Tang, Chenyi Zhou, Libin Wang +3
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on pred…
MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules
Tong Xu, Xinzhe Cao, Zhihui Zhu +2
Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of…
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10
The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…
ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question Answering
Xiaoke Guo, Songze Li, Zhiqiang Liu +4
Table serialization remains a critical bottleneck for Large Language Models (LLMs) in complex table question answering, hindered by challenges such as structural neglect, represent…
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
Ziwen Xu, Kewei Xu, Haoming Xu +8
Large Language Models (LLMs) are increasingly deployed in socially sensitive domains, yet their unpredictable behaviors, ranging from misaligned intent to inconsistent personality,…