23 papers
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
Neel Mokaria, Rishie Raj, Dheeraj Baiju +13
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
Xiaoqian Shen, Mohamed Elhoseiny
Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images and long-context videos rema…
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
Abdulwahab Felemban, Yahia Battach, Faizan Farooq Khan +10
Coral reefs are rapidly declining under anthropogenic pressures (e.g., climate change), creating an urgent need for scalable and automated monitoring. Progress in data-driven coral…
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang +2
Grounded video question answering (GVQA) aims to localize relevant temporal segments in videos and generate accurate answers to a given question; however, large video-language mode…
iMotion-LLM: Instruction-Conditioned Trajectory Generation
Abdulwahab Felemban, Nussair Hroub, Jian Ding +4
We introduce iMotion-LLM, a large language model (LLM) integrated with trajectory prediction modules for interactive motion generation. Unlike conventional approaches, it generates…
Step-by-step Layered Design Generation
Faizan Farooq Khan, K J Joseph, Koustava Goswami +2
Design generation, in its essence, is a step-by-step process where designers progressively refine and enhance their work through careful modifications. Despite this fundamental cha…