6 citations · 6 across the 4 of their papers we have counts for
4 papers
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Zhihao Xie, Junfeng Wu, Xinting Hu +2
Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure. We introduce VideoRAE, a repres…
Minkowski geometry of finite Hurwitz continued fractions
Yifei Gu, Lai Jiang
We study the Minkowski geometry of finite-level sets of Gaussian rationals defined by the lengths of their Hurwitz continued fraction expansions. For each , let be t…
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6
Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real w…
Performance of a modular ton-scale pixel-readout liquid argon time projection chamber
DUNE Collaboration, A. Abed Abud, B. Abi +1360
The Module-0 Demonstrator is a single-phase 600 kg liquid argon time projection chamber operated as a prototype for the DUNE liquid argon near detector. Based on the ArgonCube desi…