18 papers
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
Zhengyu Hu, Zheyuan Xiao, Linxin Song +10
LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the high…
CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?
Jintao Huang, Fengqing Jiang, Radha Poovendran +1
We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, an…
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, Lucas Goncalves +7
Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the…
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
Fengqing Jiang, Yichen Feng, Yuetai Li +3
The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publication loops where AI-generated research…
JobBench: Aligning Agent Work With Human Will
Yuetai Li, Yichen Feng, Zhangchen Xu +21
Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflow…
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
Shujie Han, Feng Jiang, Patrick P. C. Lee +5
Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster-wide outages. Existing checkp…