From the 1 of 14 linked papers with an AI index.
14 papers
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
Mingyang Wu, Kaituo Feng, Bohao Li +3
Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarci…
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
Xiangbo Gao, Siyuan Yang, Ping He +12
Visko Orbis 1.0 is a live model that generates long videos in real time, letting users change prompts on the fly while preserving subject, scene, and style consistency across hour‑…
When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation
Mingyan Wu, Han Yang, Omer Ben-Porat +1
Retrieval-Augmented Generation (RAG) typically assumes that external knowledge is free, but many high-quality sources are paywalled, licensed, restricted, or otherwise costly to ac…
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
Fangzhou Lin, Peiran Li, Lingyu Xu +12
Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture th…
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation
Zihao Zhu, Kuan-Ru Huang, Zhaoming Xu +6
High-resolution datasets are essential for advancing super-resolution (SR) and text-to-image (T2I) diffusion research. However, current publicly available datasets lack both the na…
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
Xiangbo Gao, Sicong Jiang, Bangya Liu +12
As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional…