4 papers
InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3
Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interr…
First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents
Syed Mahbubul Huq, Pranava Madhyastha
We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories fr…
Knowing Before Answering: Decoding Language Models for Reliable RAG
Syed Mahbubul Huq, Christopher Child, Tillman Weyde +1
In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer b…
Evaluating LLMs for Combinatorial Optimization: One-Phase and Two-Phase Heuristics for 2D Bin-Packing
Syed Mahbubul Huq, Daniel Brito, Daniel Sikar +3
This paper presents an evaluation framework for assessing Large Language Models' (LLMs) capabilities in combinatorial optimization, specifically addressing the 2D bin-packing probl…