3 papers
cs.LG2026
A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization
Prasanth YSS, Zhichen Ren, Rasa Hosseinzadeh +6
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization remains prone to collapse. We analyse this instability through…
cs.DC2025
Zero-Execution Retrieval-Augmented Configuration Tuning of Spark Applications
Raunaq Suri, Ilan Gofman, Guangwei Yu +1
Large-scale data processing is increasingly done using distributed computing frameworks like Apache Spark, which have a considerable number of configurable parameters that affect r…
cs.CL2025
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation
Satya Krishna Gorti, Ilan Gofman, Zhaoyan Liu +5
Text-to-SQL generation enables non-experts to interact with databases via natural language. Recent advances rely on large closed-source models like GPT-4 that present challenges in…