collaborators

6 papers

cs.LG2026

FIBER: A Differentially Private Optimizer with Filter-Aware Innovation Bias Correction

Duc Dm, Thao Do, Minh Son Hoang +3

Differentially private (DP) training protects individual examples by adding noise to gradients, but the injected noise interacts nontrivially with adaptive optimizers. Recent DP me…

cs.CV2026

SAT: Selective Aggregation Transformer for Image Super-Resolution

Dinh Phu Tran, Thao Do, Saad Wazir +3

Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational complexity of vanilla self-attenti…

cs.CL2026

LooComp: Leverage Leave-One-Out Strategy to Encoder-only Transformer for Efficient Query-aware Context Compression

Thao Do, Dinh Phu Tran, An Vo +2

Efficient context compression is crucial for improving the accuracy and scalability of question answering. For the efficiency of Retrieval Augmented Generation, context should be d…

cs.LG2026

Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering

Dinh Phu Tran, Jihoon Jeong, Saad Wazir +4

We present a formal problem formulation for \textit{Reliable} Audio-Visual Question Answering (-AVQA), where we prefer abstention over answering incorrectly. While rec…

cs.CV2025

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization

Son Nguyen, Giang Nguyen, Hung Dao +2

Key Information Extraction (KIE) underpins the understanding of visual documents (e.g., receipts and contracts) by extracting precise semantic content and accurately capturing spat…

cs.CL2025

Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition

Thao Do, Dinh Phu Tran, An Vo +1

Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While s…