6 papers
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
Priya Gurjar, Md Farhan Ishmam, Kenneth Marino
Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and repeated act…
TimeWarp: Evaluating Web Agents by Revisiting the Past
Md Farhan Ishmam, Kenneth Marino
The improvement of web agents on current benchmarks raises the question: Do today's agents perform just as well when the web changes? We introduce TimeWarp, a benchmark that emulat…
Contextual Breach: Assessing the Robustness of Transformer-based QA Models
Asir Saadat, Nahian Ibn Asad
Contextual question-answering models are susceptible to adversarial perturbations to input context, commonly observed in real-world scenarios. These adversarial noises are designed…
Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation
Md Tariquzzaman, Md Farhan Ishmam, Saiyma Sittul Muna +2
Sign Language (SL) enables two-way communication for the deaf and hard-of-hearing community, yet many sign languages remain under-resourced in the AI space. Sign Language Instructi…
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
Sadia Alam, Md Farhan Ishmam, Navid Hasin Alvee +3
The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fu…
Visual Robustness Benchmark for Visual Question Answering (VQA)
Md Farhan Ishmam, Ishmam Tashdeed, Talukder Asir Saadat +3
Can Visual Question Answering (VQA) systems perform just as well when deployed in the real world? Or are they susceptible to realistic corruption effects e.g. image blur, which can…