4 papers
Language Model Networks: Supervision-Efficient Learning through Dense Communication
Shiguang Wu, Yaqing Wang, Quanming Yao
Language models are increasingly used not only as standalone predictors but also as components in larger inference systems, from test-time scaling to multi-agent collaboration. We…
Asking LLMs to Verify First is Almost Free Lunch
Shiguang Wu, Quanming Yao
To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we introduce Verification-First (VF), a stra…
Self-Generative Adversarial Fine-Tuning for Large Language Models
Shiguang Wu, Yaqing Wang, Quanming Yao
Fine-tuning large language models (LLMs) for alignment typically relies on supervised fine-tuning or reinforcement learning from human feedback, both limited by the cost and scarci…
Learning to Learn with Contrastive Meta-Objective
Shiguang Wu, Yaqing Wang, Yatao Bian +1
Meta-learning enables learning systems to adapt quickly to new tasks, similar to humans. Different meta-learning approaches all work under/with the mini-batch episodic training fra…