Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Not All Synthetic Data Is Yours to Learn From
Sina Alemohammad, Li Chen, Richard G. Baraniuk +1
Can a language model improve from plain text sampled from itself, with no prompts, no teacher, no verifier, and no reward model? Yes, but only when the synthetic corpus is compatib…
cs.CL2026
Good Agentic Friends Do Not Just Give Verbal Advice: They Can Update Your Weights
Wenrui Bao, Huan Wang, Jian Wang +3
Multi-agent LLM systems usually collaborate by exchanging natural-language messages. This interface is simple and interpretable, but it forces each sender's intermediate computatio…
cs.CL2026
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Kyuyoung Kim, Kevin Wang, Yunfei Xie +7
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only fina…