2 papers
cs.LG2026
Search Self-play: Pushing the Frontier of Agent Capability without Supervision
Hongliang Lu, Yuhang Wen, Pengyu Cheng +7
Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on well-crafted task queries and cor…
cs.MM2026
Cap2Sum: Learning to Summarize Videos by Generating Captions
Cairong Zhao, Chutian Wang, Zifan Song +3
With the rapid growth of video data on the internet, video summarization is becoming a very important AI technology. However, due to the high labelling cost of video summarization,…