2 papers
cs.LG2026
Apertus LLM Family Expansion via Distillation and Quantization
Andrei Panferov, Davit Melikidze, Martin Jaggi +1
The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to s…
cs.LG2026
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
Davit Melikidze, Marian Schneider, Jessica Lam +4
Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring…