3 papers
cs.CL2026
: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
Quyet V. Do, Thinh Pham, Nguyen Nguyen +3
We study a pipeline that curates reasoning data from initial structured data for improving long-context reasoning in large language models (LLMs). Our approach, , constructs…
cs.CL2025
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar +58
Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing…
cs.CL2024
What Really is Commonsense Knowledge?
Quyet V. Do, Junze Li, Tung-Duong Vuong +3
Commonsense datasets have been well developed in Natural Language Processing, mainly through crowdsource human annotation. However, there are debates on the genuineness of commonse…