Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Anchored Supervised Fine-Tuning
He Zhu, Junyou Su, Peng Lai +4
Post-training of large language models involves a fundamental trade-off between supervised fine-tuning (SFT), which efficiently mimics demonstrations but tends to memorize, and rei…
cs.LG2024
A Critical Review of Causal Reasoning Benchmarks for Large Language Models
Linying Yang, Vik Shirvaikar, Oscar Clivio +1
Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retr…