#multimodal models

try —

15 papers match

cs.LG2026

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

Zihan Dong, Rui Qian, Qishi Zhan +3

The paper introduces Adaptive Anticipatory Policy Trees (AAPT), a method that pre‑computes conditional action trees during idle screen time so GUI agents can react instantly to eve…

#gui automation#policy trees#latency reduction#multimodal models
cs.AI2026

IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD

Nianchen Deng, Jiaxin Ai, Tao Hu +10

The paper introduces IndustryForge-27B, a multimodal foundation model fine‑tuned on diverse industrial CAD data to understand drawings, generate parametric modeling scripts, and co…

#multimodal models#industrial cad#parametric modeling#code generation
cs.CV2026

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Siyu Yan, Zhuoran Yan, Haiying Xu +10

The paper presents See2Think, an evaluation framework and benchmark for testing whether multimodal large language models actually use intermediate visual states during reasoning, a…

#multimodal models#visual reasoning#intermediate visual states#benchmark evaluation
cs.CL2026

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

Farhan Farsi, Shayan Bali, Mohammad Heydari Rad +2

The paper studies how multimodal large language models associate musical instruments with gender categories, creating a new dataset (Symphony-Bias) and finding that text modalities…

#gender bias#multimodal models#musical instruments#stereotype analysis
cs.CV2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Zheng Tong, Yang Liu, Wanshu Fan +6

The paper reviews how large language and multimodal models are being used as autonomous agents in medical tasks, covering their architectures, applications, evaluation methods, and…

#agentic ai#clinical decision support#multimodal models#evaluation
cs.CV2026

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2

The paper presents Symbal, a dual‑stage method that uses off‑the‑shelf foundation models to automatically detect systematic misalignments—recurring caption errors tied to specific…

#caption evaluation#systematic errors#multimodal models#benchmark
cond-mat.other2026

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

A. C. Opus, J. Q. Lu

The paper describes how a modern multimodal assistant model, MiniCPM-V-4.6, was fully implemented and run on an older 2011 NVIDIA Tesla C2075 GPU using all‑GPU CUDA inference, deta…

#multimodal models#gpu inference#cuda optimization#vision-language integration
cs.SE2026

When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face

Yujian Liu, Xiao Yu, Jacky Keung +3

The paper empirically examines user discussions on Hugging Face to understand how people perceive general-purpose and multimodal large language models, identifying key concerns suc…

#large language models#multimodal models#user perception#empirical study
cs.CV2026

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

Yiyang Yao, Shanglin Liu, Jianming Lv +4

The paper introduces AspectCLIP, a method that groups image-text pairs by shared textual aspects and applies consistency regularization within these groups to avoid forcing unrelat…

#contrastive learning#image-text alignment#representation learning#aspect-guided regularization
cs.CV2026

HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents

Hy Vision Team, Huawen Shen, Zhengyang Tang +20

The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…

#mobile gui agents#multimodal models#data scaling#environment simulation
cs.CV2026

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2

The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…

#long-video reasoning#evidence retrieval#temporal grounding#chain-of-thought
cs.CL2026

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task

Anand Kamble, Aniket Tathe

The paper re-implements the NAVER LABS IWSLT instruction-following system for the 2026 shared task, using SeamlessM4T-v2-large as a speech encoder and Qwen3-4B-Instruct as the LLM,…

#speech translation#instruction following#multimodal models#synthetic data generation
cs.AI2026

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Zeyu Chen, Huanjin Yao, Ziwang Zhao +1

The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…

#multimodal models#evaluation benchmarks#large language models#chain-of-thought
cs.CV2026

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Yuqi Tang, Tengfei Liu, Yizheng Lai +18

The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…

#video generation#keyframe conditioning#benchmark#evaluation metrics
cs.LG2026

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

Jing Liu, Chenxuanyin Zou, Jiayang Ren +5

The paper introduces FedCMM, a framework that combines elastic weight consolidation, synthetic replay, and task‑similarity‑aware gradient aggregation to prevent catastrophic forget…

#federated learning#continual learning#multimodal models#elastic weight consolidation

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.