activity
20172026
most citedA Deep Reinforcement Learning Chatbot

200 citations · 233 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

Hanxian Huang, Igor Fedorov, Andrey Gromov +14

Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The most useful OD-LLMs produce nea…

cs.LG2025

MobileLLM-Pro Technical Report

Patrick Huber, Ernie Chang, Wei Wen +16

Efficient on-device language models around 1 billion parameters are essential for powering low-latency AI applications on mobile and wearable devices. However, achieving strong per…

cs.LG2025

CoSMoEs: Compact Sparse Mixture of Experts

Patrick Huber, Akshat Shrivastava, Ernie Chang +3

Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mi…

cs.LG201913 cited

Neural Assistant: Joint Action Prediction, Response Generation, and Latent Knowledge Reasoning

Arvind Neelakantan, Semih Yavuz, Sharan Narang +5

Task-oriented dialog presents a difficult challenge encompassing multiple problems including multi-turn language understanding and generation, knowledge retrieval and reasoning, an…

cs.LG2019

Deep Reinforcement Learning For Modeling Chit-Chat Dialog With Discrete Attributes

Chinnadhurai Sankar, Sujith Ravi

Open domain dialog systems face the challenge of being repetitive and producing generic responses. In this paper, we demonstrate that by conditioning the response generation on int…

cs.LG2018

The Bottleneck Simulator: A Model-based Deep Reinforcement Learning Approach

Iulian Vlad Serban, Chinnadhurai Sankar, Michael Pieper +2

Deep reinforcement learning has recently shown many impressive successes. However, one major obstacle towards applying such methods to real-world problems is their lack of data-eff…