activity
20232026
most citedCones: Concept Neurons in Diffusion Models for Customized Generation

19 citations · 30 across the 10 of their papers we have counts for

collaborators

10 papers

cs.RO2026

The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents

Ziyu Wang, Chenyuan Liu, Yushun Xiang +16

Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack…

cs.CV20241 cited

Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs

Shi Liu, Kecheng Zheng, Wei Chen

Existing Large Vision-Language Models (LVLMs) primarily align image features of vision encoder with Large Language Models (LLMs) to leverage their superior text generation capabili…

cs.CV2024

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

Ziyuan Huang, Kaixiang Ji, Biao Gong +6

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence…

cs.CV2024

DreamLIP: Language-Image Pre-training with Long Captions

Kecheng Zheng, Yifei Zhang, Wei Wu +5

Language-image pre-training largely relies on how precisely and thoroughly a text describes its paired image. In practice, however, the contents of an image can be so rich that wel…

cs.CV20232 cited

AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort

Wen Wang, Canyu Zhao, Hao Chen +3

Story visualization aims to generate a series of images that match the story described in texts, and it requires the generated images to satisfy high quality, alignment with the te…

cs.CV2023

Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis

Jiapeng Zhu, Ceyuan Yang, Kecheng Zheng +3

Due to the difficulty in scaling up, generative adversarial networks (GANs) seem to be falling from grace on the task of text-conditioned image synthesis. Sparsely-activated mixtur…