4 papers · 1 filter
Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval
João Maria Janeiro, Mathurin Videau, Andrea Caciolai +3
Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically…
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood…
Structural Deep Encoding for Table Question Answering
Raphaël Mouravieff, Benjamin Piwowarski, Sylvain Lamprier
Although Transformers-based architectures excel at processing textual information, their naive adaptation for tabular data often involves flattening the table structure. This simpl…
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
Camille Barboule, Benjamin Piwowarski, Yoan Chabot
The field of visually-rich document understanding, which involves interacting with visually-rich documents (whether scanned or born-digital), is rapidly evolving and still lacks co…