3 papers
cs.CV2025
Learning Fine-to-Coarse Cuboid Shape Abstraction
Gregor Kobsik, Morten Henkel, Yanjiang He +4
The abstraction of 3D objects with simple geometric primitives like cuboids allows to infer structural information from complex geometry. It is important for 3D shape understanding…
cs.CV2024
Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data Generation
Tim Elsner, Paula Usinger, Julius Nehring-Wirxel +5
In language processing, transformers benefit greatly from text being condensed. This is achieved through a larger vocabulary that captures word fragments instead of plain character…
cs.CV2024
Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data
Tim Elsner, Paula Usinger, Victor Czech +4
In quantised autoencoders, images are usually split into local patches, each encoded by one token. This representation is redundant in the sense that the same number of tokens is s…