3 papers
cs.CL2026
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
Seyyed Saeid Cheshmi, Hahnemann Ortiz, James Mooney +1
Vision-language models (VLMs) have demonstrated strong reasoning abilities in literal multimodal tasks such as visual mathematics and science question answering. However, figurativ…
cs.LG2025
Scaling Unverifiable Rewards: A Case Study on Visual Insights
Shuyu Gan, James Mooney, Pan Hao +4
Large Language Model (LLM) agents can increasingly automate complex reasoning through Test-Time Scaling (TTS), iterative refinement guided by reward signals. However, many real-wor…
cs.LG2025
A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
Shuyu Gan, Renxiang Wang, James Mooney +1
Automating end-to-end data science pipeline with AI agents still stalls on two gaps: generating insightful, diverse visual evidence and assembling it into a coherent, professional…