2 papers
cs.CV2025
Step1X-Edit: A Practical Framework for General Image Editing
Shiyu Liu, Yucheng Han, Peng Xing +21
In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has i…
cs.CV2024
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
Zhi Cai, Songtao Liu, Guodong Wang +3
DETR has set up a simple end-to-end pipeline for object detection by formulating this task as a set prediction problem, showing promising potential. Despite its notable advancement…