activity
20212024
most citedTextPainter: Multimodal Text Image Generation with Visual-harmony and Text-comprehension for Poster Design

6 citations · 16 across the 14 of their papers we have counts for

collaborators

14 papers

cs.IR2024

Enhancing Taobao Display Advertising with Multimodal Representations: Challenges, Approaches and Insights

Xiang-Rong Sheng, Feifan Yang, Litong Gong +10

Despite the recognized potential of multimodal data to improve model accuracy, many large-scale industrial recommendation systems, including Taobao display advertising system, pred…

cs.CV2024

Enhancing Prompt Following with Visual Control Through Training-Free Mask-Guided Diffusion

Hongyu Chen, Yiqi Gao, Min Zhou +4

Recently, integrating visual controls into text-to-image~(T2I) models, such as ControlNet method, has received significant attention for finer control capabilities. While various t…

cs.CV2024

Tuning-Free Noise Rectification for High Fidelity Image-to-Video Generation

Weijie Li, Litong Gong, Yiran Zhu +4

Image-to-video (I2V) generation tasks always suffer from keeping high fidelity in the open domains. Traditional image animation techniques primarily focus on specific domains such…

cs.CV20241 cited

AtomoVideo: High Fidelity Image-to-Video Generation

Litong Gong, Yiran Zhu, Weijie Li +4

Recently, video generation has achieved significant rapid development based on superior text-to-image generation techniques. In this work, we propose a high fidelity framework for…

cs.CL20241 cited

ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models

Yanan Wu, Jie Liu, Xingyuan Bu +10

This paper introduces ConceptMath, a bilingual (English and Chinese), fine-grained benchmark that evaluates concept-wise mathematical reasoning of Large Language Models (LLMs). Unl…

cs.CL2024

E^2-LLM: Efficient and Extreme Length Extension of Large Language Models

Jiaheng Liu, Zhiqi Bai, Yuanxing Zhang +11

Typically, training LLMs with long context sizes is computationally expensive, requiring extensive training hours and GPU resources. Existing long-context extension methods usually…