4 papers
UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs
Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo +4
We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substi…
VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark
Amirhossein Dabiriaghdam, Shayan Vassef, Mohammadreza Bakhtiari +5
Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a problem through a tool and then re…
When Minor Edits Matter: LLM-Driven Prompt Attack for Medical VLM Robustness in Ultrasound
Yasamin Medghalchi, Milad Yazdani, Amirhossein Dabiriaghdam +7
Ultrasound is widely used in clinical practice due to its portability, cost-effectiveness, safety, and real-time imaging capabilities. However, image acquisition and interpretation…
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
Amirhossein Dabiriaghdam, Lele Wang
The widespread adoption of large language models (LLMs) necessitates reliable methods to detect LLM-generated text. We introduce SimMark, a robust sentence-level watermarking algor…