activity
20242026
most citedA Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

2 citations · 2 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2026

Seek to Segment: Active Perception for Panoramic Referring Segmentation

Song Tang, Shuming Hu, Xincheng Shuai +2

Existing referring segmentation models passively process static images captured from fixed perspectives, limiting their applicability in Embodied AI, where agents must perform acti…

cs.CV2026

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

Jinyu Liu, Xincheng Shuai, Henghui Ding +1

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically asse…

cs.AI2026

Training the Knowledge Base through Evidence Distillation and Write-Back Enrichment

Yuxing Lu, Xukai Zhao, Wei Wu +1

The knowledge base in a retrieval-augmented generation (RAG) system is typically assembled once and never revised, even though the facts a query requires are often fragmented acros…

cs.CV2026

GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

Xincheng Shuai, Ziye Li, Henghui Ding +1

Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality s…

cs.CV2025

SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation

Zhenyuan Qin, Xincheng Shuai, Henghui Ding

Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simulta…

cs.CV2025

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine

Xincheng Shuai, Zhenyuan Qin, Henghui Ding +1

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation.…