3 citations · 4 across the 8 of their papers we have counts for
8 papers
BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data
Xuwu Wang, Qiwen Cui, Yunzhe Tao +16
Large language models (LLMs) have become increasingly pivotal across various domains, especially in handling complex data types. This includes structured data processing, as exempl…
ViTAR: Vision Transformer with Any Resolution
Qihang Fan, Quanzeng You, Xiaotian Han +5
This paper tackles a significant challenge faced by Vision Transformers (ViTs): their constrained scalability across different image resolutions. Typically, ViTs experience a perfo…
-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model
Yufeng Zhang, Liyu Chen, Boyi Liu +4
Recent advances in reinforcement learning (RL) algorithms aim to enhance the performance of language models at scale. Yet, there is a noticeable absence of a cost-effective and sta…
InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding
Haogeng Liu, Quanzeng You, Xiaotian Han +7
Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intri…
Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling
Haogeng Liu, Qihang Fan, Tingkai Liu +5
This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…
DeVAn: Dense Video Annotation for Video-Language Models
Tingkai Liu, Yunzhe Tao, Haogeng Liu +5
We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeV…