activity
20222024
most citedGemma: Open Models Based on Gemini Research and Technology

238 citations · 293 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024238 cited

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105

This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate stro…

cs.CL2023

Triggering Multi-Hop Reasoning for Question Answering in Language Models using Soft Prompts and Random Walks

Kanishka Misra, Cicero Nogueira dos Santos, Siamak Shakeri

Despite readily memorizing world knowledge about entities, pre-trained language models (LMs) struggle to compose together two or more facts to perform multi-hop reasoning in questi…

cs.CV202339 cited

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…

cs.CV20232 cited

Generative Models for 3D Point Clouds

Lingjie Kong, Pankaj Rajak, Siamak Shakeri

Point clouds are rich geometric data structures, where their three dimensional structure offers an excellent domain for understanding the representation learning and generative mod…

cs.CL20238 cited

Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models

Renat Aksitov, Chung-Ching Chang, David Reitter +2

Despite recent progress, it has been difficult to prevent semantic hallucinations in generative Large Language Models. One common solution to this is augmenting LLMs with a retriev…

cs.CL20226 cited

Reducing Retraining by Recycling Parameter-Efficient Prompts

Brian Lester, Joshua Yurtsever, Siamak Shakeri +1

Parameter-efficient methods are able to use a single frozen pre-trained large language model (LLM) to perform many tasks by learning task-specific soft prompts that modulate model…