2 papers
cs.CL2024
Byte Latent Transformer: Patches Scale Better Than Tokens
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11
We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant imp…
cs.CL2024
ALMA: Alignment with Minimal Annotation
Michihiro Yasunaga, Leonid Shamis, Chunting Zhou +4
Recent approaches to large language model (LLM) alignment typically require millions of human annotations or rely on external aligned models for synthetic data generation. This pap…