4 papers
BenchBench: Benchmarking Automated Benchmark Generation
Yandan Zheng, Haoran Luo, Zhenghong Lin +2
Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are c…
OrchMAS: Orchestrated Reasoning with Multi Collaborative Heterogeneous Scientific Expert Structured Agents
Yichao Feng, Haoran Luo, Zhenghong Lin +4
Multi-agent large language model frameworks are promising for complex multi step reasoning, yet existing systems remain weak for scientific and knowledge intensive domains due to s…
Probing Large Language Models in Reasoning and Translating Complex Linguistic Puzzles
Zheng-Lin Lin, Yu-Fei Shih, Shu-Kai Hsieh
This paper investigates the utilization of Large Language Models (LLMs) for solving complex linguistic puzzles, a domain requiring advanced reasoning and adept translation capabili…
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
Yu-Fei Shih, Zheng-Lin Lin, Shu-Kai Hsieh
We explore the capabilities of LVLMs and LLMs in deciphering rare scripts not encoded in Unicode. We introduce a novel approach to construct a multimodal dataset of linguistic puzz…