24 papers
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
Mizanur Rahman, Arshia Azimlu, Shadikur Rahman +4
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is…
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as…
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
Gundeep Singh, Parsa Kavehzadeh, Jing Xia +5
Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional business intelligence tools…
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar +3
Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios,…
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
Md Tahmid Rahman Laskar, Xue-Yong Fu, Seyyed Saeed Sarfjoo +3
Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be conve…
Assessing the Quality of Mental Health Support in LLM Responses through Multi-Attribute Human Evaluation
Abeer Badawi, Md Tahmid Rahman Laskar, Elahe Rahimi +6
The escalating global mental health crisis, marked by persistent treatment gaps, availability, and a shortage of qualified therapists, positions Large Language Models (LLMs) as a p…