6 papers · 1 filter
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
Raoyuan Zhao, Yihong Liu, Yupei Du +2
Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evaluation and training pipelines…
GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis
Yu Du, Wenlong Zhu, Xingze Li +3
Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained structural reasoning. We intro…
Disentangling the Roles of Representation and Selection in Data Pruning
Yupei Du, Yingjin Song, Hugh Mee Wong +3
Data pruning, selecting small but impactful subsets, offers a promising way to efficiently scale NLP model training. However, existing methods often involve many different design c…
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
Yingjin Song, Yupei Du, Denis Paperno +1
This paper introduces the TempVS benchmark, which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models (MLLMs) in image sequences. TempVS co…
FTFT: Efficient and Robust Fine-Tuning by Transferring Training Dynamics
Yupei Du, Albert Gatt, Dong Nguyen
Despite the massive success of fine-tuning Pre-trained Language Models (PLMs), they remain susceptible to out-of-distribution input. Dataset cartography is a simple yet effective d…
Transforming Dutch: Debiasing Dutch Coreference Resolution Systems for Non-binary Pronouns
Goya van Boven, Yupei Du, Dong Nguyen
Gender-neutral pronouns are increasingly being introduced across Western languages. Recent evaluations have however demonstrated that English NLP systems are unable to correctly pr…