4 papers
LLM Benchmark Datasets Should Be Contamination-Resistant
Ali Al-Lawati, Jason Lucas, Dongwon Lee +1
Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretr…
Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue
Ali Al-Lawati, Nafis Tripto, Abolfazl Ansari +3
The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that a…
Graph-based Molecular In-context Learning Grounded on Morgan Fingerprints
Ali Al-Lawati, Jason Lucas, Zhiwei Zhang +2
In-context learning (ICL) effectively conditions large language models (LLMs) for molecular tasks, such as property prediction and molecule captioning, by embedding carefully selec…
Semantic Captioning: Benchmark Dataset and Graph-Aware Few-Shot In-Context Learning for SQL2Text
Ali Al-Lawati, Jason Lucas, Prasenjit Mitra
Large Language Models (LLMs) have demonstrated remarkable performance in various NLP tasks, including semantic parsing, which translates natural language into formal code represent…