Publications (10)
Experimental Assortments for Choice Estimation and Nest Identification
Xintong Yu, Will Ma, Michael Zhao
What assortments (subsets of items) should be offered, to collect data for estimating a choice model over total items? We propose a structured, non-adaptive experiment design r…
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
Minbin Huang, Han Shi, Chuanyang Zheng +5
Modern Mixture-of-Experts (MoE) architectures allocate expert capacity through a rigid per-layer rule: each transformer layer owns a separate expert set. This convention couples de…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
METGEN: A Module-Based Entailment Tree Generation Framework for Answer Explanation
Ruixin Hong, Hongming Zhang, Xintong Yu +1
Knowing the reasoning chains from knowledge to the predicted answers can help construct an explainable question answering (QA) system. Advances on QA explanation propose to explain…
Exophoric Pronoun Resolution in Dialogues with Topic Regularization
Xintong Yu, Hongming Zhang, Yangqiu Song +3
Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem. Previous works on pronoun coreference resolution (PCR) mostly f…
VD-PCR: Improving Visual Dialog with Pronoun Coreference Resolution
Xintong Yu, Hongming Zhang, Ruixin Hong +2
The visual dialog task requires an AI agent to interact with humans in multi-round dialogs based on a visual environment. As a common linguistic phenomenon, pronouns are often used…
What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues
Xintong Yu, Hongming Zhang, Yangqiu Song +2
Grounding a pronoun to a visual object it refers to requires complex reasoning from various information sources, especially in conversational scenarios. For example, when people in…
Efficient Text-Guided 3D-Aware Portrait Generation with Score Distillation Sampling on Distribution
Yiji Cheng, Fei Yin, Xiaoke Huang +5
Text-to-3D is an emerging task that allows users to create 3D content with infinite possibilities. Existing works tackle the problem by optimizing a 3D representation with guidance…
Foreground segmentation based on multi-resolution and matting
Xintong Yu, Xiaohan Liu, Yisong Chen
We propose a foreground segmentation algorithm that does foreground extraction under different scales and refines the result by matting. First, the input image is filtered and resa…
ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts
Zhida Feng, Zhenyu Zhang, Xintong Yu +12
Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution im…