10 papers
VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval
Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng +1
Large language models are increasingly used in practical systems, making efficient model selection important for reducing deployment cost. LLM routing has emerged as a practical so…
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
Yow-Fu Liou, Yu-Chien Tang, Yu-Hsiang Liu +1
Benchmarking large language models (LLMs) is critical for understanding their capabilities, limitations, and robustness. In addition to interface artifacts, prior studies have show…
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark
Zheqing Li, Yiying Yang, Jiping Lang +16
Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks primarily depend on exam-style or…
ConceptKT: A Benchmark for Concept-Level Deficiency Prediction in Knowledge Tracing
Yu-Chen Kang, Yu-Chien Tang, An-Zi Yen
Knowledge Tracing (KT) is a critical technique for modeling student knowledge to support personalized learning. However, most KT systems focus on binary correctness prediction and…
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
Wei-Ling Hsu, Yu-Chien Tang, An-Zi Yen
The increasing reliance on Large Language Models (LLMs) across various domains extends to education, where students progressively use generative AI as a tool for learning. While pr…
CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models
Yong-En Tian, Yu-Chien Tang, An-Zi Yen +1
Aspect-based summarization has attracted significant attention for its ability to generate more fine-grained and user-aligned summaries. While most existing approaches assume a set…