2 papers
cs.CL2026
ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions
Peixian Zhou, Yuxu Chen, Chaorui Zhang +3
Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogi…
cs.CL2025
AutoMix: Automatically Mixing Language Models
Pranjal Aggarwal, Aman Madaan, Ankit Anand +10
Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively le…