3 papers
cs.CV2026
Token-Efficient Multimodal Reasoning via Image Prompt Packaging
Joong Ho Choi, Jiayang Zhao, Avani Appalla +3
Deploying large multimodal language models at scale is constrained by token-based inference costs, yet the cost-performance behavior of visual prompting strategies remains poorly c…
cs.AI2025
CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
Joong Ho Choi, Jiayang Zhao, Jeel Shah +5
Large Language Models (LLMs) deliver powerful reasoning and generation capabilities but incur substantial run-time costs when operating in agentic workflows that chain together len…
cs.AI2025
A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain
William Flanagan, Mukunda Das, Rajitha Ramanayake +10
As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical mach…