activity
20242026
collaborators

6 papers

cs.AI2026

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

Masaharu Mizumoto, Dat Nguyen, Zhiheng Han +6

Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarit…

cs.CL2025

Improving Table Understanding with LLMs and Entity-Oriented Search

Thi-Nhung Nguyen, Hoang Ngo, Dinh Phung +2

Our work addresses the challenges of understanding tables. Existing methods often struggle with the unpredictable nature of table content, leading to a reliance on preprocessing an…

cs.CL2025

Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding

Thi-Nhung Nguyen, Hoang Ngo, Dinh Phung +2

Table understanding is key to addressing challenging downstream tasks such as table-based question answering and fact verification. Recent works have focused on leveraging Chain-of…

cs.CL2025

VM14K: First Vietnamese Medical Benchmark

Thong Nguyen, Duc Nguyen, Minh Dang +6

Medical benchmarks are indispensable for evaluating the capabilities of language models in healthcare for non-English-speaking communities,therefore help ensuring the quality of re…

cs.CL2025

ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations

Quang Hieu Pham, Thuy Duong Nguyen, Tung Pham +2

The capabilities of large language models (LLMs) have been enhanced by training on data that reflects human thought processes, such as the Chain-of-Thought format. However, evidenc…

cs.CL2024

Who's Who: Large Language Models Meet Knowledge Conflicts in Practice

Quang Hieu Pham, Hoang Ngo, Anh Tuan Luu +1

Retrieval-augmented generation (RAG) methods are viable solutions for addressing the static memory limits of pre-trained language models. Nevertheless, encountering conflicting sou…