◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Xinlei Chen

5 papers hereh-index 8982 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4
  • last author1

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CV3
  • cs.LG1
  • cs.RO1
same name
  • Xinlei Chen — 19 papers, h 7
  • Xinlei Chen — 8 papers, h 3
  • Xinlei Chen — 8 papers, h 6
  • Xinlei Chen — 6 papers, h 3
  • Xinlei Chen — 6 papers, h 4
  • Xinlei Chen — 5 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.CV2025

Meta CLIP 2: A Worldwide Scaling Recipe

Yung-Sung Chuang, Yang Li, Dong Wang +13

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (M…

cs.LG2025

Transformers without Normalization

Jiachen Zhu, Xinlei Chen, Kaiming He +2

Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve th…

cs.CV2025

Scaling Language-Free Visual Representation Learning

David Fan, Shengbang Tong, Jiachen Zhu +8

Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This mul…

cs.RO2025

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression

Lirui Wang, Kevin Zhao, Chaoqi Liu +1

We propose Heterogeneous Masked Autoregression (HMA) for modeling action-video dynamics to generate high-quality data and evaluation in scaling robot learning. Building interactive…

cs.CV2024

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Shengbang Tong, David Fan, Jiachen Zhu +7

In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.