Publications (25)
Strong quantum computational advantage using a superconducting quantum processor
Yulin Wu, Wan-Su Bao, Sirui Cao +51
Scaling up to a large number of qubits with high-precision control is essential in the demonstrations of quantum computational advantage to exponentially outpace the classical hard…
Precision Enhancement of 3D Surfaces from Multiple Compressed Depth Maps
Pengfei Wan, Gene Cheung, Philip A. Chou +3
In texture-plus-depth representation of a 3D scene, depth maps from different camera viewpoints are typically lossily compressed via the classical transform coding / coefficient qu…
Multimodal active speaker detection and virtual cinematography for video conferencing
Ross Cutler, Ramin Mehran, Sam Johnson +4
Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zoom…
Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution
Emad Barsoum, Cha Zhang, Cristian Canton Ferrer +1
Crowd sourcing has become a widely adopted scheme to collect ground truth labels. However, it is a well-known problem that these labels can be very noisy. In this paper, we demonst…
From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding
Li Sun, Florian Luisier, Kayhan Batmanghelich +2
Current state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. This process known as tokenization relies…
KOSMOS-2.5: A Multimodal Literate Model
Tengchao Lv, Yupan Huang, Jingye Chen +13
The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI). In this paper we present KOSMOS-2.5, a m…