most citedVehicle-Infrastructure Cooperative 3D Object Detection via Feature Flow Prediction

14 citations · 49 across the 8 of their papers we have counts for

collaborators

13 papers

cs.CV2024

PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Junsong Chen, Chongjian Ge, Enze Xie +7

In this paper, we introduce PixArt-Σ, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-Σrepresents a significant advancement over…

cs.CV20241 cited

TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model

Jiahao Lyu, Jin Wei, Gangyan Zeng +4

Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene…

cs.CV20241 cited

Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Zhenyu Wang, Enze Xie, Aoxue Li +3

Despite significant advancements in text-to-image models for generating high-quality images, these methods still struggle to ensure the controllability of text prompts over images…

cs.AI20243 cited

A Survey of Reasoning with Foundation Models

Jiankai Sun, Chuanyang Zheng, Enze Xie +31

Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It…

cs.CV20241 cited

PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models

Junsong Chen, Yue Wu, Simian Luo +5

This technical report introduces PIXART-δ, a text-to-image synthesis framework that integrates the Latent Consistency Model (LCM) and ControlNet into the advanced PIXART-α model. P…

cs.CV202310 cited

Flow-Based Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection

Haibao Yu, Yingjuan Tang, Enze Xie +3

Cooperatively utilizing both ego-vehicle and infrastructure sensor data can significantly enhance autonomous driving perception abilities. However, the uncertain temporal asynchron…