activity
20242026
collaborators

6 papers

cs.CV2026

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

Mingyang Wu, Kaituo Feng, Bohao Li +3

Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarci…

cs.CV2026

Twins: Learn to Predict Unified Representations with Focal Loss

Kaixiong Gong, Xin Cai, Bin Lin +9

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codeb…

cs.CL2026

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Zhekai Chen, Chengqi Duan, Kaiyue Sun +4

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assist…

cs.CV2025

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li +7

Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically expl…

cs.CV2024

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Kaixiong Gong, Kaituo Feng, Bohao Li +8

Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini 1.5 Pro, and Reka Core, have expanded their capabilities to include vision and audio modalities. While th…

cs.CV2024

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Sijie Cheng, Kechen Fang, Yangyang Yu +6

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoTh…