2 citations · 2 across the 7 of their papers we have counts for
9 papers
Seek to Segment: Active Perception for Panoramic Referring Segmentation
Song Tang, Shuming Hu, Xincheng Shuai +2
Existing referring segmentation models passively process static images captured from fixed perspectives, limiting their applicability in Embodied AI, where agents must perform acti…
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation
Jinyu Liu, Xincheng Shuai, Henghui Ding +1
Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically asse…
Training the Knowledge Base through Evidence Distillation and Write-Back Enrichment
Yuxing Lu, Xukai Zhao, Wei Wu +1
The knowledge base in a retrieval-augmented generation (RAG) system is typically assembled once and never revised, even though the facts a query requires are often fragmented acros…
GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
Xincheng Shuai, Ziye Li, Henghui Ding +1
Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality s…
SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
Zhenyuan Qin, Xincheng Shuai, Henghui Ding
Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simulta…
Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
Xincheng Shuai, Zhenyuan Qin, Henghui Ding +1
Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation.…