activity
20222025
most citedMatryoshka Query Transformer for Large Vision-Language Models

1 citations · 2 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan +24

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from…

cs.CV20241 cited

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

Zi-Yi Dou, Xitong Yang, Tushar Nagarajan +5

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-s…

cs.CV20241 cited

Matryoshka Query Transformer for Large Vision-Language Models

Wenbo Hu, Zi-Yi Dou, Liunian Harold Li +3

Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong…

cs.CV2023

ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos

Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu +5

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and languag…

cs.CV2023

Masked Path Modeling for Vision-and-Language Navigation

Zi-Yi Dou, Feng Gao, Nanyun Peng

Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited a…