7 papers
Punching Above Their Weight: Classification-Head Fine-Tuning of Tiny Language Models (TLMs) for Verifiable Multiple-Choice Tasks
Bhavesh Sood, Jaromir Savelka
We define Tiny Language Models (TLMs) as models below roughly 3B parameters that fit on mainstream consumer devices. We study how to adapt them for and use them on verifiable multi…
Towards Spec Learning: Inference-Time Alignment from Preference Pairs
Dhriti Krishnan, Tejas Goyal, Jaromir Savelka
Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's resp…
mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?
Yerzhan Sapenov, Jaromir Savelka
We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consis…
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
Li Zhang, Jaromir Savelka, Kevin Ashley
Multi-label legal annotation requires assigning multiple labels from large, evolving taxonomies to long, fact-intensive documents, often under limited supervision. Parametric encod…
MCQ Difficulty Prediction via Modeling Learner Heterogeneity Using Data-Driven Cognitive Profiling
Dhriti Krishnan, Jaromir Savelka
Predicting the difficulty of multiple-choice questions (MCQs) is important for effective assessment, yet current methods typically assume a unimodal student ability distribution, o…
Changes in Coding Behavior and Performance Since the Introduction of LLMs
Yufan Zhang, Jaromir Savelka, Seth Copen Goldstein +1
The widespread availability of large language models (LLMs) has changed how students engage with coding and problem-solving. While these tools may increase student productivity, th…