6 papers
LADBench: A Benchmark for Logical Fault Detection in Images
Sahasra Kondapalli, Lara Radovanovic, Aadi Palnitkar +2
Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning remains underexplored. Existing a…
FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning
Mingyang Mao, Bhargav Rishi Medisetti, Utkarsh Grover +4
Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health…
Embodied Foundation Models at the Edge: A Survey of Deployment Constraints and Mitigation Strategies
Utkarsh Grover, Ravi Ranjan, Mingyang Mao +9
Deploying foundation models in embodied edge systems is fundamentally a systems problem, not just a problem of model compression. Real-time control must operate within strict size,…
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
Aadi Palnitkar, Mingyang Mao, Nicholas Waytowich +2
As large language models (LLMs) are applied to increasingly longer and more complex tasks, there is a growing need for realistic long-context benchmarks that require selective read…
Adaptive Dynamics Planning for Robot Navigation
Yuanjie Lu, Mingyang Mao, Tong Xu +3
Autonomous robot navigation systems often rely on hierarchical planning, where global planners compute collision-free paths without considering dynamics, and local planners enforce…
Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding
Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri +3
To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent age…