3 papers
cs.LG2026
Prompt replay: speeding up grpo with on-policy reuse of high-signal prompts
Andrei Baroian, Rutger Berger
Reinforcement learning with verifiable rewards (RLVR) plays a crucial role in expanding the capacities of LLM reasoning, but GRPO-style training is dominated by expensive rollouts…
cs.CL2025
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
Andrei Baroian
We study clinical Named Entity Recognition (NER) on the CADEC corpus and compare three families of approaches: (i) BERT-style encoders (BERT Base, BioClinicalBERT, RoBERTa-large),…
cs.CL2025
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
Andrei Baroian, Kasper Notebomer
Transformer-based language models traditionally use uniform (isotropic) layer sizes, yet they ignore the diverse functional roles that different depths can play and their computati…