activity
20242026
collaborators

7 papers

cs.CL2026

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak +4

While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sover…

cs.AI2026

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong +45

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedic…

cs.CV2025

What Makes for Good Image Captions?

Delong Chen, Samuel Cahyawijaya, Etsuko Ishii +3

This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode sema…

cs.CL2025

High-Dimensional Interlingual Representations of Large Language Models

Bryan Wilie, Samuel Cahyawijaya, Junxian He +1

Large language models (LLMs) trained on massive multilingual datasets hint at the formation of interlingual constructs--a shared subspace in the representation space. However, evid…

cs.CL2025

High-Dimension Human Value Representation in Large Language Models

Samuel Cahyawijaya, Delong Chen, Yejin Bang +5

The widespread application of LLMs across various tasks and fields has necessitated the alignment of these models with human values and preferences. Given various approaches of hum…

cs.CV2025

Subobject-level Image Tokenization

Delong Chen, Samuel Cahyawijaya, Jianfeng Liu +2

Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we in…