#multimodal models
15 papers match
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
Zihan Dong, Rui Qian, Qishi Zhan +3
The paper introduces Adaptive Anticipatory Policy Trees (AAPT), a method that pre‑computes conditional action trees during idle screen time so GUI agents can react instantly to eve…
IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD
Nianchen Deng, Jiaxin Ai, Tao Hu +10
The paper introduces IndustryForge-27B, a multimodal foundation model fine‑tuned on diverse industrial CAD data to understand drawings, generate parametric modeling scripts, and co…
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Siyu Yan, Zhuoran Yan, Haiying Xu +10
The paper presents See2Think, an evaluation framework and benchmark for testing whether multimodal large language models actually use intermediate visual states during reasoning, a…
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
Farhan Farsi, Shayan Bali, Mohammad Heydari Rad +2
The paper studies how multimodal large language models associate musical instruments with gender categories, creating a new dataset (Symphony-Bias) and finding that text modalities…
Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
Zheng Tong, Yang Liu, Wanshu Fan +6
The paper reviews how large language and multimodal models are being used as autonomous agents in medical tasks, covering their architectures, applications, evaluation methods, and…
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2
The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…
A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
A. C. Opus, J. Q. Lu
The paper describes how a modern multimodal assistant model, MiniCPM-V-4.6, was fully implemented and run on an older 2011 NVIDIA Tesla C2075 GPU using all‑GPU CUDA inference, deta…
When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
Yujian Liu, Xiao Yu, Jacky Keung +3
The paper empirically examines user discussions on Hugging Face to understand how people perceive general-purpose and multimodal large language models, identifying key concerns suc…
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
Yiyang Yao, Shanglin Liu, Jianming Lv +4
The paper introduces AspectCLIP, a method that groups image-text pairs by shared textual aspects and applies consistency regularization within these groups to avoid forcing unrelat…
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…
NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task
Anand Kamble, Aniket Tathe
The paper re-implements the NAVER LABS IWSLT instruction-following system for the 2026 shared task, using SeamlessM4T-v2-large as a speech encoder and Qwen3-4B-Instruct as the LLM,…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Yuqi Tang, Tengfei Liu, Yizheng Lai +18
The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
Jing Liu, Chenxuanyin Zou, Jiayang Ren +5
The paper introduces FedCMM, a framework that combines elastic weight consolidation, synthetic replay, and task‑similarity‑aware gradient aggregation to prevent catastrophic forget…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.