Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
Houston H. Zhang, Tao Zhang, Li Gu +5
Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipel…
cs.CV2026
ASMIL: Attention-Stabilized Multiple Instance Learning for Whole Slide Imaging
Linfeng Ye, Shayan Mohajer Hamidi, Zhixiang Chi +5
Attention-based multiple instance learning (MIL) has emerged as a powerful framework for whole slide image (WSI) diagnosis, leveraging attention to aggregate instance-level feature…
cs.CV2025
Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
Houston H. Zhang, Tao Zhang, Baoze Lin +10
User interface to code (UI2Code) aims to generate executable code that can faithfully reconstruct a given input UI. Prior work focuses largely on web pages and mobile screens, leav…