16 papers
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support
Mizanur Rahman, Abeer Badawi, Elahe Rahimi +4
Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive…
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar +3
Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios,…
DATAREEL: Automated Data-Driven Video Story Generation with Animations
Ridwan Mahbub, Syem Aziz, Mahir Ahmed +4
Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communicati…
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
Mizanur Rahman, Mohammed Saidul Islam, Md Tahmid Rahman Laskar +2
Text-to-Visualization (Text2Vis) systems translate natural language queries over tabular data into concise answers and executable visualizations. While closed-source LLMs generate…
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
Ahmed Masry, Megh Thakkar, Patrice Bechard +9
Retrieval-augmented generation has proven practical when models require specialized knowledge or access to the latest data. However, existing methods for multimodal document retrie…
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang +19
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps vi…