54 citations · 89 across the 6 of their papers we have counts for
7 papers · 1 filter
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Lijie Fan, Luming Tang, Siyang Qin +11
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture p…
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
Benlin Liu, Yuhao Dong, Yiqin Wang +6
Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. C…
RealFill: Reference-Driven Generation for Authentic Image Completion
Luming Tang, Nataniel Ruiz, Qinghao Chu +8
Recent advances in generative imagery have brought forth outpainting and inpainting models that can produce high-quality, plausible image content in unknown regions. However, the c…
Emergent Correspondence from Image Diffusion
Luming Tang, Menglin Jia, Qianqian Wang +2
Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explici…
Few-Shot Classification with Feature Map Reconstruction Networks
Davis Wertheimer, Luming Tang, Bharath Hariharan
In this paper we reformulate few-shot classification as a reconstruction problem in latent space. The ability of the network to reconstruct a query feature map from support feature…
Revisiting Pose-Normalization for Fine-Grained Few-Shot Recognition
Luming Tang, Davis Wertheimer, Bharath Hariharan
Few-shot, fine-grained classification requires a model to learn subtle, fine-grained distinctions between different classes (e.g., birds) based on a few images alone. This requires…