4 papers
Regularize or Localize: When Training-Time KV-Cache Geometry Pays Under Quantization
Libo Sun, Po-Wei Harn, Zewei Zhang +2
We study whether \sigreg -- LeJEPA's anti-collapse objective -- can reshape representations during standard autoregressive language-model pretraining, and when the resulting geomet…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Miracle Kang, Lights Shi, Lucy Liang +10
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…
Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks
Libo Sun, Po-Wei Harn, Zewei Zhang +2
Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{retrieval-warmed energy-based r…
4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on Relational DBs
Minjie Wang, Quan Gan, David Wipf +17
Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably…