3 papers
cs.CL2026
CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding
Wei Jia, Zhicong Lu, Yu Chen +6
Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods pr…
cs.CL2026
CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment
Jiayue Cao, Zhicong Lu, Xuehan Sun +6
Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenario…
cs.AI2025
A Vision for Auto Research with LLM Agents
Chengwei Liu, Chong Wang, Jiayue Cao +16
This paper introduces Agent-Based Auto Research, a structured multi-agent framework designed to automate, coordinate, and optimize the full lifecycle of scientific research. Levera…