most citedViTAR: Vision Transformer with Any Resolution

3 citations · 4 across the 8 of their papers we have counts for

collaborators

8 papers

cs.AI2024

BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Xuwu Wang, Qiwen Cui, Yunzhe Tao +16

Large language models (LLMs) have become increasingly pivotal across various domains, especially in handling complex data types. This includes structured data processing, as exempl…

cs.CV2024★ 3 cited

ViTAR: Vision Transformer with Any Resolution

Qihang Fan, Quanzeng You, Xiaotian Han +5

This paper tackles a significant challenge faced by Vision Transformers (ViTs): their constrained scalability across different image resolutions. Typically, ViTs experience a perfo…

cs.LG2024

-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model

Yufeng Zhang, Liyu Chen, Boyi Liu +4

Recent advances in reinforcement learning (RL) algorithms aim to enhance the performance of language models at scale. Yet, there is a noticeable absence of a cost-effective and sta…

cs.CV2024

InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Haogeng Liu, Quanzeng You, Xiaotian Han +7

Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intri…

cs.CV2023★ 1 cited

Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Haogeng Liu, Qihang Fan, Tingkai Liu +5

This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…

cs.CV2023

DeVAn: Dense Video Annotation for Video-Language Models

Tingkai Liu, Yunzhe Tao, Haogeng Liu +5

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeV…