3 papers
cs.LG2026
Autoregressive Synthesis of Sparse and Semi-Structured Mixed-Type Data
Thomas RückstieÃ, Robin Vujanic
Synthetic data generation is an important capability for privacy-preserving data sharing, system benchmarking and test data provisioning. For mixed-type data, existing synthesizers…
cs.IR2026
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
Robin Vujanic, Thomas Rueckstiess
We present LEAF ("Lightweight Embedding Alignment Framework"), a knowledge distillation framework for text embedding models. A key distinguishing feature is that our distilled leaf…
cs.LG2024
ORIGAMI: A generative transformer architecture for predictions from semi-structured data
Thomas RückstieÃ, Alana Huang, Robin Vujanic
Despite the popularity and widespread use of semi-structured data formats such as JSON, end-to-end supervised learning applied directly to such data remains underexplored. We prese…