papers

Publications (26)

cs.MM2021

Grid-VLP: Revisiting Grid Features for Vision-Language Pre-training

Ming Yan, Haiyang Xu, Chenliang Li +4

Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…

cs.IR2025

Markup Language Modeling for Web Document Understanding

Su Liu, Bin Bi, Jan Bakus +3

Web information extraction (WIE) is an important part of many e-commerce systems, supporting tasks like customer analysis and product recommendation. In this work, we look at the p…

cs.CL2019

Incorporating External Knowledge into Machine Reading for Generative Question Answering

Bin Bi, Chen Wu, Ming Yan +3

Commonsense and background knowledge is required for a QA model to answer many nontrivial questions. Different from existing work on knowledge-aware QA, we focus on a more challeng…

cs.CL2026

Reinforcement Learning for LLM Post-Training: A Survey

Zhichao Wang, Kiran Ramnath, Bin Bi +9

Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle in domains like math and coding.…

cs.CL2017

A Neural Comprehensive Ranker (NCR) for Open-Domain Question Answering

Bin Bi, Hao Ma

This paper proposes a novel neural machine reading model for open-domain question answering at scale. Existing machine comprehension models typically assume that a short piece of r…

cs.CV2023

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Haiyang Xu, Qinghao Ye, Ming Yan +12

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…