69 citations · 116 across the 8 of their papers we have counts for
11 papers · 1 filter
Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Munan Ning, Bin Zhu, Yujia Xie +5
Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user i…
Album Storytelling with Iterative Story-aware Captioning and Large Language Models
Munan Ning, Yujia Xie, Dongdong Chen +5
This work studies how to transform an album to vivid and coherent stories, a task we refer to as "album storytelling". While this task can help preserve memories and facilitate exp…
Image as First-Order Norm+Linear Autoregression: Unveiling Mathematical Invariance
Yinpeng Chen, Xiyang Dai, Dongdong Chen +4
This paper introduces a novel mathematical property applicable to diverse images, referred to as FINOLA (First-Order Norm+Linear Autoregressive). FINOLA represents each image in th…
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Junke Wang, Dongdong Chen, Chong Luo +4
Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…
Self-Supervised Learning based on Heat Equation
Yinpeng Chen, Xiyang Dai, Dongdong Chen +4
This paper presents a new perspective of self-supervised learning based on extending heat equation into high dimensional feature space. In particular, we remove time dependence by…
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based v…