Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
Jui-Ming Yao, Bing-Cheng Xie, Sheng-Wei Peng +5
Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlo…
cs.AI2025
Verbal Process Supervision Elicits Better Coding Agents
Hao-Yuan Chen, Cheng-Pong Huang, Jui-Ming Yao
The emergence of large language models and their applications as AI agents have significantly advanced state-of-the-art code generation benchmarks, transforming modern software eng…