1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7
We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…
cs.CV2024★ 1 cited
End-to-end multi-modal product matching in fashion e-commerce
Sándor Tóth, Stephen Wilson, Alexia Tsoukara +3
Product matching, the task of identifying different representations of the same product for better discoverability, curation, and pricing, is a key capability for online marketplac…