3 papers
cs.CL2026
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
Aryo Pradipta Gema, Neel Rajani, Rohit Saxena +2
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bia…
cs.LG2025
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
Neel Rajani, Aryo Pradipta Gema, Seraphina Goldfarb-Tarrant +1
Training large language models (LLMs) for reasoning via maths and code datasets has become a major new focus in LLM post-training. Two particularly popular approaches are reinforce…
cs.CL2024
KodeXv0.1: A Family of State-of-the-Art Financial Large Language Models
Neel Rajani, Lilli Kiessling, Aleksandr Ogaltsov +1
Although powerful, current cutting-edge LLMs may not fulfil the needs of highly specialised sectors. We introduce KodeXv0.1, a family of large language models that outclass GPT-4 i…