most citedText-Video Multi-Grained Integration for Video Moment Montage

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2025

EEA: Exploration-Exploitation Agent for Long Video Understanding

Te Yang, Xiangyu Zhu, Bo Wang +3

Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding…

cs.CV20241 cited

Text-Video Multi-Grained Integration for Video Moment Montage

Zhihui Yin, Ye Ma, Xipeng Cao +3

The proliferation of online short video platforms has driven a surge in user demand for short video editing. However, manually selecting, cropping, and assembling raw footage into…

cs.CV2024

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

Zhentao Tan, Ben Xue, Jian Jia +7

This paper presents the \textbf{S}emantic-a\textbf{W}ar\textbf{E} spatial-t\textbf{E}mporal \textbf{T}okenizer (SweetTok), a novel video tokenizer to overcome the limitations in cu…

cs.CV2024

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Siqi Kou, Jiachun Jin, Zhihong Liu +6

We introduce Orthus, an autoregressive (AR) transformer that excels in generating images given textual prompts, answering questions based on visual inputs, and even crafting length…

cs.CV2024

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Te Yang, Jian Jia, Xiangyu Zhu +9

Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…