5 papers
TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval
Xinyu Sun, Huangyu Dai, Lingtao Mao +5
E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval se…
UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
Xinyu Nan, Lingtao Mao, Huangyu Dai +8
Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current…
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
Zexin Zheng, Huangyu Dai, Lingtao Mao +8
Traditional vision search, similar to search and recommendation systems, follows the multi-stage cascading architecture (MCA) paradigm to balance efficiency and conversion. Specifi…
Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation
Guohong Huang, Ling-An Zeng, Zexin Zheng +2
We propose a novel approach for generating text-guided human-object interactions (HOIs) that achieves explicit joint-level interaction modeling in a computationally efficient manne…
iManip: Skill-Incremental Learning for Robotic Manipulation
Zexin Zheng, Jia-Feng Cai, Xiao-Ming Wu +3
The development of a generalist agent with adaptive multiple manipulation skills has been a long-standing goal in the robotics community. In this paper, we explore a crucial task,…