papers

Publications (25)

quant-ph2021

Strong quantum computational advantage using a superconducting quantum processor

Yulin Wu, Wan-Su Bao, Sirui Cao +51

Scaling up to a large number of qubits with high-precision control is essential in the demonstrations of quantum computational advantage to exponentially outpace the classical hard…

cs.CV2014

Precision Enhancement of 3D Surfaces from Multiple Compressed Depth Maps

Pengfei Wan, Gene Cheung, Philip A. Chou +3

In texture-plus-depth representation of a 3D scene, depth maps from different camera viewpoints are typically lossily compressed via the classical transform coding / coefficient qu…

eess.AS2022

Multimodal active speaker detection and virtual cinematography for video conferencing

Ross Cutler, Ramin Mehran, Sam Johnson +4

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zoom…

cs.CV2016

Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution

Emad Barsoum, Cha Zhang, Cristian Canton Ferrer +1

Crowd sourcing has become a widely adopted scheme to collect ground truth labels. However, it is a well-known problem that these labels can be very noisy. In this paper, we demonst…

cs.CL2023

From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding

Li Sun, Florian Luisier, Kayhan Batmanghelich +2

Current state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. This process known as tokenization relies…

cs.CL2024

KOSMOS-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +13

The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI). In this paper we present KOSMOS-2.5, a m…