13 papers
Scaling an Autoregressive Transformer for Single-Cell Generation
Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog
We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors o…
m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning
Yosub Shin, Michael Buriek, Igor Molybog
Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead repres…
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
Amirhossein Bozorgkhoo, Igor Molybog
Speculative decoding is a technique that uses multiple language models to accelerate infer- ence. Previous works have used an experi- mental approach to optimize the throughput of…
Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols
Yosub Shin, Michael Buriek, Boris Sobolev +5
We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constan…
SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
Amila Indika, Igor Molybog
Numerous knowledge workers utilize spreadsheets in business, accounting, and finance. However, a lack of systematic documentation methods for spreadsheets hinders automation, colla…
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Seokhee Hong, Sunkyoung Kim, Guijin Son +3
The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applica…