4 citations · 7 across the 17 of their papers we have counts for
21 papers
SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation
Tong Bao, Mir Tafseer Nayeem, Yi Zhao +2
Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXi…
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar +3
Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios,…
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
Mir Tafseer Nayeem, Davood Rafiei
Large language models (LLMs) are increasingly deployed in high-stakes domains, yet they expose only limited language settings, most notably "English (US)," despite the global diver…
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
Sawsan Alqahtani, Mir Tafseer Nayeem, Md Tahmid Rahman Laskar +2
Tokenization underlies every large language model, yet it remains an under-theorized and inconsistently designed component. Common subword approaches such as Byte Pair Encoding (BP…
OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
Mir Tafseer Nayeem, Davood Rafiei
We study the problem of opinion highlights generation from large volumes of user reviews, often exceeding thousands per entity, where existing methods either fail to scale or produ…
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
Mir Tafseer Nayeem, Sawsan Alqahtani, Md Tahmid Rahman Laskar +2
Tokenization is a crucial but under-evaluated step in large language models (LLMs). The standard metric, fertility (the average number of tokens per word), captures compression eff…