works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
most citedFORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets

1 citations · 1 across the 4 of their papers we have counts for

collaborators

11 papers

cs.HC2026

Surprise2Refine: Axis-Centered Exploration-To-Refinement for Agent-Assisted Creative Scaffolding

Yuzhe You, Gromit Yeuk-Yin Chan, Shunan Guo +4

Designers require different design spaces across creative stages: broad during exploration, and targeted during refinement. Yet existing agent-driven tools assume a fixed or contin…

cs.IR2026

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping

Jiacheng Chen, Tao Zhang, Manxi Lin +26

ShopX is a foundation model that directly translates user shopping intents into item-space actions using semantic IDs, integrating intent understanding, planning, and item retrieva…

cs.CV2026

MMAgent-R: Learning to Rerank and Reject for Agentic mRAG

Tao Zhang, Ziqi Zhang, Zongyang Ma +7

Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…

cs.IR20261 cited

FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets

Kairui Fu, Tao Zhang, Shuwen Xiao +9

Semantic identifiers (SIDs) have gained increasing attention in generative retrieval (GR) for recommendation due to their meaningful semantic discriminability. However, current stu…

cs.CV2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

Bob Zhang, Haoran Li, Tao Zhang +5

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instruct…

cs.CV2026

Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs

Houston H. Zhang, Tao Zhang, Baoze Lin +10

User interface to code (UI2Code) aims to generate executable code that can faithfully reconstruct a given input UI. Prior work focuses largely on web pages and mobile screens, leav…