2 papers
cs.CV2025
KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
Son Hai Nguyen, Diwei Wang, Jinhyeok Jang +1
Accurate vision-based action recognition is crucial for developing autonomous robots that can operate safely and reliably in complex, real-world environments. In this work, we adva…
cs.CV2025
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
Son Nguyen, Giang Nguyen, Hung Dao +2
Key Information Extraction (KIE) underpins the understanding of visual documents (e.g., receipts and contracts) by extracting precise semantic content and accurately capturing spat…