activity
20232026
most citedGPT-4V(ision) as A Social Media Analysis Engine

4 citations · 10 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Jianhui Wei, Jie Tan, Hengchuan Zhu +6

Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimo…

cs.CV2026

How Far Are Video Models from True Multimodal Reasoning?

Xiaotian Zhang, Jianhui Wei, Yuan Wang +9

Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing…

cs.CV2026

A Versatile Multimodal Agent for Multimedia Content Generation

Daoan Zhang, Wenlin Yao, Xiaoyang Wang +3

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and…

cs.CV2026

VisualActBench: Can VLMs See and Act like a Human?

Daoan Zhang, Pai Liu, Xiaofei Zhou +6

Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely…

cs.CV2026

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

Kai Liu, Jungang Li, Yuchong Sun +13

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-…

cs.CV2025

LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models

Shuai Wang, Daoan Zhang, Tianyi Bai +3

Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-…