activity
20242026
collaborators

7 papers

cs.CL2026

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…

cs.CL2026

Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs

Guy Mor-Lan, Omer Goldman, Matan Eyal +6

Multilingual large language models (LLMs) have minimized the fluency gap between languages. This advancement, however, exposes models to the risk of biased behavior, as knowledge a…

cs.CV2026

FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images

Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…

cs.CL2025

ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer

Omer Goldman, Uri Shaham, Dan Malkin +11

To achieve equitable performance across languages, large language models (LLMs) must be able to abstract knowledge beyond the language in which it was learnt. However, the current…

cs.LG2025

Assay2Mol: large language model-based drug design using BioAssay context

Yifan Deng, Spencer S. Ericksen, Anthony Gitter

Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional res…

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision underst…