Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Compared to What? Baselines and Metrics for Counterfactual Prompting
Zihao Yang, Mosh Levy, Yoav Goldberg +1
Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithfulness. But in this work we ar…
cs.CL2025
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
Chi Hang, Ruiqi Deng, Lavender Yao Jiang +4
Clinical measurements such as blood pressures and respiration rates are critical in diagnosing and monitoring patient outcomes. It is an important component of biomedical data, whi…
cs.CL2024
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
Yanbing Chen, Ruilin Wang, Zihao Yang +2
Packing and shuffling tokens is a common practice in training auto-regressive language models (LMs) to prevent overfitting and improve efficiency. Typically documents are concatena…