papers

Publications (6)

eess.AS2023

2-bit Conformer quantization for automatic speech recognition

Oleg Rybakov, Phoenix Meadowlark, Shaojin Ding +4

Large speech models are rapidly gaining traction in research community. As a result, model compression has become an important topic, so that these models can fit in memory and be…

cs.CV2016

Improving Facial Analysis and Performance Driven Animation through Disentangling Identity and Expression

David Rim, Sina Honari, Md Kamrul Hasan +1

We present techniques for improving performance driven facial animation, emotion recognition, and facial key-point or landmark prediction using learned identity invariant represent…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

eess.AS2024

USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models

Shaojin Ding, David Qiu, David Rim +10

End-to-end automatic speech recognition (ASR) models have seen revolutionary quality gains with the recent development of large-scale universal speech models (USM). However, deploy…

eess.AS2023

RAND: Robustness Aware Norm Decay For Quantized Seq2seq Models

David Qiu, David Rim, Shaojin Ding +2

With the rapid increase in the size of neural networks, model compression has become an important area of research. Quantization is an effective technique at decreasing the model s…

cs.CL2026

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…