activity
20222024
most citedEmu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

30 citations · 42 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV20244 cited

Comprehensive Performance Evaluation of YOLOv11, YOLOv10, YOLOv9, YOLOv8 and YOLOv5 on Object Detection of Power Equipment

Zijian He, Kang Wang, Tian Fang +3

With the rapid development of global industrial production, the demand for reliability in power equipment has been continuously increasing. Ensuring the stability of power system o…

cs.DC20241 cited

Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Vikranth Srivatsa, Zijian He, Reyna Abhyankar +2

Prompts to large language models (LLMs) have evolved beyond simple user questions. For LLMs to solve complex problems, today's practices are to include domain-specific instructions…

cs.CV2024

AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment

Yuanfeng Xu, Yuhao Chen, Zhongzhan Huang +4

Recent video editing advancements rely on accurate pose sequences to animate subjects. However, these efforts are not suitable for cross-species animation due to pose misalignment…

cs.RO2024

Legged Robot State Estimation within Non-inertial Environments

Zijian He, Sangli Teng, Tzu-Yuan Lin +2

This paper investigates the robot state estimation problem within a non-inertial environment. The proposed state estimation approach relaxes the common assumption of static ground…

cs.CV20243 cited

Cache Me if You Can: Accelerating Diffusion Models through Block Caching

Felix Wimbauer, Bichen Wu, Edgar Schoenfeld +11

Diffusion models have recently revolutionized the field of image synthesis due to their ability to generate photorealistic images. However, one of the major drawbacks of diffusion…

cs.CV202330 cited

Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Xiaoliang Dai, Ji Hou, Chih-Yao Ma +23

Training text-to-image models with web scale image-text pairs enables the generation of a wide range of visual concepts from text. However, these pre-trained models often face chal…