works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

Guoxuan Chen, Chufeng Xiao, Haoran Yang +30

Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…

cs.CV2026

AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models

Jintao Lin, Bowen Dong, Weikang Shi +4

The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering…

cs.CV2025

FIRM: Flexible Interactive Reflection reMoval

Xiao Chen, Xudong Jiang, Yunkang Tao +4

Removing reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfacto…

cs.CV2024

SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality

Chenyang Lei, Liyi Chen, Jun Cen +5

Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many d…

cs.CV2024

SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality

Chenyang Lei, Liyi Chen, Jun Cen +6

Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many d…