2 papers
cs.CL2026
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Jiajun Shi, Siyuan Tao, Yuhao Wu +18
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them…
cs.CV2026
Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG
Zhe Jin, Zhimin Lin, Bin Zheng +2
Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from…