1 citations · 1 across the 3 of their papers we have counts for
5 papers
Pitfalls of Evaluating Language Models with Open Benchmarks
Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar +3
Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility…
Benchmarking LLMs on the Semantic Overlap Summarization Task
John Salvador, Naman Bansal, Mousumi Akter +3
Semantic Overlap Summarization (SOS) is a constrained multi-document summarization task, where the constraint is to capture the common/overlapping information between two alternati…
LLMs as On-demand Customizable Service
Souvika Sarkar, Mohammad Fakhruddin Babar, Monowar Hasan +1
Large Language Models (LLMs) have demonstrated remarkable language understanding and generation capabilities. However, training, deploying, and accessing these models pose notable…
Introducing "Forecast Utterance" for Conversational Data Science
Md Mahadi Hassan, Alex Knipper, Shubhra Kanti Karmaker
Envision an intelligent agent capable of assisting users in conducting forecasting tasks through intuitive, natural conversations, without requiring in-depth knowledge of the under…
The Daunting Dilemma with Sentence Encoders: Success on Standard Benchmarks, Failure in Capturing Basic Semantic Properties
Yash Mahajan, Naman Bansal, Shubhra Kanti Karmaker
In this paper, we adopted a retrospective approach to examine and compare five existing popular sentence encoders, i.e., Sentence-BERT, Universal Sentence Encoder (USE), LASER, Inf…