activity
20242026
collaborators

11 papers

cs.CV2026

Visual Prompt Discovery via Semantic Exploration

Jaechang Kim, Yotaro Shimose, Zhao Wang +3

LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, which incorporate image manipulation co…

cs.AI2025

WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation

Kuang-Da Wang, Zhao Wang, Yotaro Shimose +2

Witnessed by the recent advancements on leveraging LLM for coding and multimodal understanding, we present WebGen-V, a new benchmark and framework for instruction-to-HTML generatio…

cs.CV2025

BannerAgency: Advertising Banner Design with Multimodal LLM Agents

Heng Wang, Yotaro Shimose, Shingo Takamatsu

Advertising banners are critical for capturing user attention and enhancing advertising campaign effectiveness. Creating aesthetically pleasing banner designs while conveying the c…

cs.IR2025

Forecasting Clicks in Digital Advertising: Multimodal Inputs and Interpretable Outputs

Briti Gangopadhyay, Zhao Wang, Shingo Takamatsu

Forecasting click volume is a key task in digital advertising, influencing both revenue and campaign strategy. Traditional time series models rely solely on numerical data, often o…

cs.CV2025

DesignLab: Designing Slides Through Iterative Detection and Correction

Jooyeol Yun, Heng Wang, Yotaro Shimose +2

Designing high-quality presentation slides can be challenging for non-experts due to the complexity involved in navigating various design choices. Numerous automated tools can sugg…

cs.CV2025

Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents

Zhao Wang, Bowen Chen, Yotaro Shimose +3

Recent generative models such as GPT-4o have shown strong capabilities in producing high-quality images with accurate text rendering. However, commercial design tasks like advertis…