activity
20242026
collaborators

14 papers

cs.RO2026

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue

Xingyao Lin, Xinghao Zhu, Tianyi Lu +6

Embodied agents are intelligent systems designed to perceive, reason, and act within the physical world. While the robotics community has long strived to build such versatile agent…

cs.SE2026

Written by AI, Managed by AI: Semantic Space Control and Index Sickness Elimination Across 391 Consecutive Sessions

Hui Zhang, Shuren Song

The prevailing engineering intuition for addressing conceptual drift in long-horizon LLM collaboration is to trade more formal constraints for more reliable outputs -- designing sy…

cs.RO2026

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

Tianyi Lu, Hui Zhang, Zijie Diao +8

Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To add…

cs.CV2026

NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

Xin Li, Yeying Jin, Suhang Yao +95

This paper presents an overview of the NTIRE 2026 Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images. Building upon the success of the first edition, this c…

cs.CV2026

FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data

Peng Yuan, Bingyin Mei, Hui Zhang

Composed Image Retrieval (CIR) retrieves target images using a reference image paired with modification text. Despite rapid advances, all existing methods and datasets operate at t…

cs.CV2026

CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

Hui Zhang, Dexiang Hong, Maoke Yang +7

Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation u…