3 papers
cs.LG2026
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
Haoxiang Jiang, Zihan Dong, Tianci Liu +5
Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address th…
cs.CL2025
TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription
Guo Yutong, Wanying Wang, Yue Wu +2
Table Visual Question Answering (Table VQA) is typically addressed by large vision-language models (VLMs). While such models can answer directly from images, they often miss fine-g…
cs.AI2025
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
Wanying Wang, Zeyu Ma, Xuhong Wang +3
As Large Language Models (LLMs) are increasingly deployed in highly specialized vertical domains, the evaluation of their domain-specific performance becomes critical. However, exi…