activity
20202026
most citedUPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance

11 citations · 21 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CL2026★ 2 cited

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CL2025

TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning

Mingyu Zheng, Zhifan Feng, Jia Wang +4

Despite the commendable progress of recent LLM-based data synthesis methods, they face two limitations in generating table instruction tuning data. First, they can not thoroughly e…

cs.CL2024

Textualized Agent-Style Reasoning for Complex Tasks by Multiple Round LLM Generation

Chen Liang, Zhifan Feng, Zihe Liu +4

Chain-of-thought prompting significantly boosts the reasoning ability of large language models but still faces three issues: hallucination problem, restricted interpretability, and…

cs.CV2023

Improving Video Retrieval by Adaptive Margin

Feng He, Qi Wang, Zhifan Feng +4

Video retrieval is becoming increasingly important owing to the rapid emergence of videos on the Internet. The dominant paradigm for video retrieval learns video-text representatio…

cs.CV2022★ 1 cited

CLOP: Video-and-Language Pre-Training with Knowledge Regularizations

Guohao Li, Hu Yang, Feng He +4

Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner,…

cs.CV2022★ 11 cited

UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance

Wei Li, Xue Xu, Xinyan Xiao +8

Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusio…