1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.SD2025
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
Yanhao Jia, Ji Xie, S Jivaganesh +3
Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reli…
cs.CV2025
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Zechuan Zhang, Ji Xie, Yu Lu +2
Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive d…
cs.CV2025★ 1 cited
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
Dewei Zhou, Ji Xie, Zongxin Yang +1
The growing demand for controllable outputs in text-to-image generation has driven significant advancements in multi-instance generation (MIG), enabling users to define both instan…