6 papers
An Information-Theoretic Perspective on LLM Tokenizers
Mete Erdogan, Abhiram Gorle, Shubham Chandak +2
Large language model (LLM) tokenizers act as structured compressors: by mapping text to discrete token sequences, they determine token count (and thus compute and context usage) an…
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
Yasmine Omri, Connor Ding, Tsachy Weissman +1
Modern vision language pipelines are driven by RGB vision encoders trained on massive image text corpora. While these pipelines have enabled impressive zero-shot capabilities and s…
Information-computation trade-offs in non-linear transforms
Connor Ding, Abhiram Rao Gorle, Jiwon Jeong +2
In this work, we explore the interplay between information and computation in non-linear transform-based compression for broad classes of modern information-processing tasks. We fi…
Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs
Hao Kang, Qingru Zhang, Han Cai +4
Large language models (LLMs) have shown remarkable performance across diverse reasoning and generation tasks, and are increasingly deployed as agents in dynamic environments such a…
LZMidi: Compression-Based Symbolic Music Generation
Connor Ding, Abhiram Gorle, Sagnik Bhattacharya +3
Recent advances in symbolic music generation primarily rely on deep learning models such as Transformers, GANs, and diffusion models. While these approaches achieve high-quality re…
Universal Discrete Filtering with Lookahead or Delay
Pumiao Yan, Jiwon Jeong, Naomi Sagan +1
We consider the universal discrete filtering problem, where an input sequence generated by an unknown source passes through a discrete memoryless channel, and the goal is to estima…