1 citations · 2 across the 4 of their papers we have counts for
4 papers · 1 filter
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
Jie Zhu, Hanghang Ma, Jia Wang +6
In this work, we introduce Wallaroo, a simple autoregressive baseline that leverages next-token prediction to unify multi-modal understanding, image generation, and editing at the…
LongCat-Image Technical Report
Meituan LongCat Team, Hanghang Ma, Haoxian Tan +10
We introduce LongCat-Image, a pioneering open-source and bilingual (Chinese-English) foundation model for image generation, designed to address core challenges in multilingual text…
LinVT: Empower Your Image-level Large Language Model to Understand Videos
Lishuai Gao, Yujie Zhong, Yingsen Zeng +3
Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a modu…
Video Temporal Relationship Mining for Data-Efficient Person Re-identification
Siyu Chen, Dengjie Li, Lishuai Gao +3
This paper is a technical report to our submission to the ICCV 2021 VIPriors Re-identification Challenge. In order to make full use of the visual inductive priors of the data, we t…