collaborators

13 papers

cs.LG2026

Scaling an Autoregressive Transformer for Single-Cell Generation

Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog

We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors o…

cs.CV2026

m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning

Yosub Shin, Michael Buriek, Igor Molybog

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead repres…

cs.CL2026

Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple

Amirhossein Bozorgkhoo, Igor Molybog

Speculative decoding is a technique that uses multiple language models to accelerate infer- ence. Previous works have used an experi- mental approach to optimize the throughput of…

cs.AI2026

Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols

Yosub Shin, Michael Buriek, Boris Sobolev +5

We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constan…

cs.SE2025

SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations

Amila Indika, Igor Molybog

Numerous knowledge workers utilize spreadsheets in business, accounting, and finance. However, a lack of systematic documentation methods for spreadsheets hinders automation, colla…

cs.CL2025

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

Seokhee Hong, Sunkyoung Kim, Guijin Son +3

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applica…