10 papers
LOKI: Memory-Free Null-Space Constrained Lifelong Knowledge Editing
Masih Eskandar, Miquel Sirera Perelló, Stratis Ioannidis +1
Lifelong knowledge editing aims to efficiently and sequentially update language models over time, as new knowledge becomes available or when the model makes mistakes, while preserv…
CD-RCM: Generalizable Continuous-Depth Novel View Synthesis for Reflectance Confocal Microscopy
Tooba Imtiaz, Milind Rajadhyaksha, Kivanc Kose +1
Reflectance confocal microscopy (RCM) provides noninvasive, cellular-resolution "optical biopsies" of human skin \emph{in vivo} by acquiring en-face images at successive depths, fo…
ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
Arash Akbari, Arman Akbari, Masih Eskandar +11
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressiv…
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
Le Jiang, Xiangyu Bai, Bishoy Galoaa +7
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360 video from a single image and a caption. Existing panoramic video methods optimi…
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Juyi Lin, Arash Akbari, Yumei He +13
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, ev…
LVT: Large-Scale Scene Reconstruction via Local View Transformers
Tooba Imtiaz, Lucy Chai, Kathryn Heal +4
Large transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer's well-known quadratic complexity makes it diff…