papers

Publications (15)

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2025

AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization

Junjie Shentu, Matthew Watson, Noura Al Moubayed

Text-to-image (T2I) customization empowers users to adapt the T2I diffusion model to new concepts absent in the pre-training dataset. On this basis, capturing multiple new concepts…

cond-mat.mtrl-sci2022

One-dimensional electronic states in a natural misfit structure

Alla Chikina, Gargee Bhattacharyya, Davide Curcio +7

Misfit compounds are thermodynamically stable stacks of two-dimensional materials, forming a three-dimensional structure that remains incommensurate in one direction parallel to th…

cs.LG2026

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

Yunting Song, Matthew Watson, Peter Grabowski +1

The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumula…

cs.CV2024

Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation

Junjie Shentu, Matthew Watson, Noura Al Moubayed

Subject-driven text-to-image diffusion models empower users to tailor the model to new concepts absent in the pre-training dataset using a few sample images. However, prevalent sub…

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision underst…