4 papers
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
Hoang Phan, Xianjun Yang, Yuanshun Yao +6
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for c…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data
Deren Lei, Yaxi Li, Siyao Li +6
Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI)…
InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance
Rui Xu, Mengya Hu, Deren Lei +6
The proliferation of AI-generated images has intensified the need for robust content authentication methods. We present InvisMark, a novel watermarking technique designed for high-…