1 citations · 1 across the 9 of their papers we have counts for
4 papers · 1 filter
DAPD: Dual-Anchored Policy Distillation
Jianyu Wu, Yizhou Wang, Encheng Su +2
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege ill…
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
Lintao Wang, Encheng Su, Jiaqi Liu +11
Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. E…
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Encheng Su, Jianyu Wu, Chen Tang +9
As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of…
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
Yiheng Wang, Yixin Chen, Shuo Li +33
We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…