activity
20242026
collaborators

8 papers

eess.AS2026

MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model

The Hieu Pham, Tan Dat Nguyen, Phuong Thanh Tran +2

Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that adv…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

q-bio.BM2025

A Geometric Graph-Based Deep Learning Model for Drug-Target Affinity Prediction

Md Masud Rana, Farjana Tasnim Mukta, Duc D. Nguyen

In structure-based drug design, accurately estimating the binding affinity between a candidate ligand and its protein receptor is a central challenge. Recent advances in artificial…

cs.CV2025

MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation

Kiet Dang Vu, Trung Thai Tran, Duc Dung Nguyen

Precisely localizing 3D objects from a single image constitutes a central challenge in monocular 3D detection. While DETR-like architectures offer a powerful paradigm, their direct…

eess.AS2025

Wanna hear your voice? A sample is all we need!

The Hieu Pham, Phuong Thanh Tran Nguyen, Xuan Tho Nguyen +2

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. Ho…

cs.CV2025

Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance

Duc-Hai Pham, Duc-Dung Nguyen, Anh Pham +4

Accurate prediction of 3D semantic occupancy from 2D visual images is vital in enabling autonomous agents to comprehend their surroundings for planning and navigation. State-of-the…