papers

Publications (10)

stat.ME2026

Experimental Assortments for Choice Estimation and Nest Identification

Xintong Yu, Will Ma, Michael Zhao

What assortments (subsets of items) should be offered, to collect data for estimating a choice model over total items? We propose a structured, non-adaptive experiment design r…

cs.LG2026

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts

Minbin Huang, Han Shi, Chuanyang Zheng +5

Modern Mixture-of-Experts (MoE) architectures allocate expert capacity through a rigid per-layer rule: each transformer layer owns a separate expert set. This convention couples de…

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CL2022

METGEN: A Module-Based Entailment Tree Generation Framework for Answer Explanation

Ruixin Hong, Hongming Zhang, Xintong Yu +1

Knowing the reasoning chains from knowledge to the predicted answers can help construct an explainable question answering (QA) system. Advances on QA explanation propose to explain…

cs.CL2021

Exophoric Pronoun Resolution in Dialogues with Topic Regularization

Xintong Yu, Hongming Zhang, Yangqiu Song +3

Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem. Previous works on pronoun coreference resolution (PCR) mostly f…

cs.CL2022

VD-PCR: Improving Visual Dialog with Pronoun Coreference Resolution

Xintong Yu, Hongming Zhang, Ruixin Hong +2

The visual dialog task requires an AI agent to interact with humans in multi-round dialogs based on a visual environment. As a common linguistic phenomenon, pronouns are often used…

cs.CL2019

What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues

Xintong Yu, Hongming Zhang, Yangqiu Song +2

Grounding a pronoun to a visual object it refers to requires complex reasoning from various information sources, especially in conversational scenarios. For example, when people in…

cs.CV2023

Efficient Text-Guided 3D-Aware Portrait Generation with Score Distillation Sampling on Distribution

Yiji Cheng, Fei Yin, Xiaoke Huang +5

Text-to-3D is an emerging task that allows users to create 3D content with infinite possibilities. Existing works tackle the problem by optimizing a 3D representation with guidance…

cs.CV2014

Foreground segmentation based on multi-resolution and matting

Xintong Yu, Xiaohan Liu, Yisong Chen

We propose a foreground segmentation algorithm that does foreground extraction under different scales and refines the result by matting. First, the input image is filtered and resa…

cs.CV2023

ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts

Zhida Feng, Zhenyu Zhang, Xintong Yu +12

Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution im…