9 papers
Visual Prompt Discovery via Semantic Exploration
Jaechang Kim, Yotaro Shimose, Zhao Wang +3
LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, which incorporate image manipulation co…
WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation
Kuang-Da Wang, Zhao Wang, Yotaro Shimose +2
Witnessed by the recent advancements on leveraging LLM for coding and multimodal understanding, we present WebGen-V, a new benchmark and framework for instruction-to-HTML generatio…
Forecasting Clicks in Digital Advertising: Multimodal Inputs and Interpretable Outputs
Briti Gangopadhyay, Zhao Wang, Shingo Takamatsu
Forecasting click volume is a key task in digital advertising, influencing both revenue and campaign strategy. Traditional time series models rely solely on numerical data, often o…
Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents
Zhao Wang, Bowen Chen, Yotaro Shimose +3
Recent generative models such as GPT-4o have shown strong capabilities in producing high-quality images with accurate text rendering. However, commercial design tasks like advertis…
OMS: On-the-fly, Multi-Objective, Self-Reflective Ad Keyword Generation via LLM Agent
Bowen Chen, Zhao Wang, Shingo Takamatsu
Keyword decision in Sponsored Search Advertising is critical to the success of ad campaigns. While LLM-based methods offer automated keyword generation, they face three major limit…
Auto-bidding in real-time auctions via Oracle Imitation Learning (OIL)
Alberto Silvio Chiappa, Briti Gangopadhyay, Zhao Wang +1
Online advertising has become one of the most successful business models of the internet era. Impression opportunities are typically allocated through real-time auctions, where adv…