1 citations · 2 across the 6 of their papers we have counts for
5 papers · 1 filter
Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data
Xuanming Zhang, Shwan Ashrafi, Aziza Mirsaidova +5
We study the reasoning behavior of large language models (LLMs) under limited computation budgets. In such settings, producing useful partial solutions quickly is often more practi…
Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments
Li Siyan, Zhen Xu, Vethavikashini Chithrra Raghuram +3
Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teachi…
VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation
Kun Qian, Shunji Wan, Claudia Tang +4
As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-tra…
DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and Rewriting
Xuanming Zhang, Anthony Diaz, Zixun Chen +4
Coherence in writing, an aspect that second-language (L2) English learners often struggle with, is crucial in assessing L2 English writing. Existing automated writing evaluation sy…
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
Xuanming Zhang, Zixun Chen, Zhou Yu
Lexical Substitution discovers appropriate substitutes for a given target word in a context sentence. However, the task fails to consider substitutes that are of equal or higher pr…