works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CV2026

Let RGB Be the Language of Vision

Timing Yang, Jinrui Yang, Xinlong Li +11

The paper proposes a unified vision framework that encodes all visual signals—including images, masks, and depth maps—as RGB images, turning diverse tasks into a common RGB-to-RGB…

cs.AI2026

Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory

Jiaqi Liu, Zipeng Ling, Shi Qiu +9

AI agents increasingly operate over extended time horizons, yet their ability to retain, organize, and recall multimodal experiences remains a critical bottleneck. Building effecti…

eess.IV2026

OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation

Letian Zhang, Sucheng Ren, Yanqing Liu +9

This paper presents a family of advanced vision encoder, named OpenVision 3, that learns a single, unified visual representation that can serve both image understanding and image g…

cs.CV2025

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Yanqing Liu, Xianhang Li, Letian Zhang +4

This paper provides a simplification on OpenVision's architecture and loss design for enhancing its training efficiency. Following the prior vision-language pretraining works CapPa…

cs.CV2025

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning

Xianhang Li, Yanqing Liu, Haoqin Tu +2

OpenAI's CLIP, released in early 2021, have long been the go-to choice of vision encoder for building multimodal foundation models. Although recent alternatives such as SigLIP have…

cs.CV2024

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Yanqing Liu, Xianhang Li, Zeyu Wang +2

Previous works show that noisy, web-crawled image-text pairs may limit vision-language pretraining like CLIP and propose learning with synthetic captions as a promising alternative…