9 papers
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen +4
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods su…
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
Jihao Qiu, Lingxi Xie, Xinyue Huo +2
This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, an active, reasoning-equipped mul…
AceTone: Bridging Words and Colors for Conditional Image Grading
Tianren Ma, Mingxiang Liao, Xijin Zhang +1
Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative…
ReDDiT: Rehashing Noise for Discrete Visual Generation
Tianren Ma, Xiaosong Zhang, Boyu Yang +2
In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous…
Building Vision Models upon Heat Conduction
Zhaozhi Wang, Yue Liu, Yunjie Tian +3
Visual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study,…
Adaptive Keyframe Sampling for Long Video Understanding
Xi Tang, Jihao Qiu, Lingxi Xie +3
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…