5 papers
SAT: Selective Aggregation Transformer for Image Super-Resolution
Dinh Phu Tran, Thao Do, Saad Wazir +3
Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational complexity of vanilla self-attenti…
LooComp: Leverage Leave-One-Out Strategy to Encoder-only Transformer for Efficient Query-aware Context Compression
Thao Do, Dinh Phu Tran, An Vo +2
Efficient context compression is crucial for improving the accuracy and scalability of question answering. For the efficiency of Retrieval Augmented Generation, context should be d…
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
Dinh Phu Tran, Dao Duy Hung, Daeyoung Kim
Video super-resolution remains a major challenge in low-level vision tasks. To date, CNN- and Transformer-based methods have delivered impressive results. However, CNNs are limited…
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
Thao Do, Dinh Phu Tran, An Vo +1
Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While s…
Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution
Dinh Phu Tran, Dao Duy Hung, Daeyoung Kim
Recently, window-based attention methods have shown great potential for computer vision tasks, particularly in Single Image Super-Resolution (SISR). However, it may fall short in c…