238 citations · 241 across the 8 of their papers we have counts for
8 papers
A Controlled Study on Long Context Extension and Generalization in LLMs
Yi Lu, Jing Nathan Yan, Songlin Yang +6
Broad textual understanding and in-context learning require language models that utilize full document contexts. Due to the implementation challenges associated with directly train…
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team, Thomas Mesnard, Cassidy Hardin +105
This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate stro…
RakutenAI-7B: Extending Large Language Models for Japanese
Rakuten Group, Aaron Levine, Connie Huang +27
We introduce RakutenAI-7B, a suite of Japanese-oriented large language models that achieve the best performance on the Japanese LM Harness benchmarks among the open 7B models. Alon…
Asking More Informative Questions for Grounded Retrieval
Sedrick Keh, Justin T. Chiu, Daniel Fried
When a model is trying to gather information in an interactive setting, it benefits from asking informative questions. However, in the case of a grounded multi-turn image identific…
Symbolic Planning and Code Generation for Grounded Dialogue
Justin T. Chiu, Wenting Zhao, Derek Chen +3
Large language models (LLMs) excel at processing and generating both text and code. However, LLMs have had limited applicability in grounded task-oriented dialogue as they are diff…
Multi-output Headed Ensembles for Product Item Classification
Hotaka Shiokawa, Pradipto Das, Arthur Toth +1
In this paper, we revisit the problem of product item classification for large-scale e-commerce catalogs. The taxonomy of e-commerce catalogs consists of thousands of genres to whi…