5 papers
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
Zihan Lin, Songhe Deng, Shuwei He +6
Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions im…
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
Zihang Lin, Huaiyuan Qin, Muli Yang +1
Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplet…
T-Stars-Poster: A Framework for Product-Centric Advertising Image Design
Hongyu Chen, Min Zhou, Jing Jiang +6
Creating advertising images is often a labor-intensive and time-consuming process. Can we automatically generate such images using basic product information like a product foregrou…
PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
Yifan Gao, Zihang Lin, Chuanbin Liu +4
Product posters, which integrate subject, scene, and text, are crucial promotional tools for attracting customers. Creating such posters using modern image generation methods is va…
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
Shang Chai, Zihang Lin, Min Zhou +3
Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has…