most citedFrom Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

2 citations

7 papers

cond-mat.mtrl-sci2026

High-speed and high-gain graphene photovoltaic phototransistor gated by a van der Waals heterojunction

Yihan Yin, Jiayi Zhang, Xiaolong Zhang +5

Two-dimensional (2D) material-based phototransistors offer a unique combination of optical sensing, signal amplification, and logic operation within a single device, yet fundamenta…

cs.CV2026

Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

Axi Niu, Jieheng Li, Kang Zhang +3

Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observ…

cs.RO2026

STAGE: STyle-controllable Action GEneration for personalized autonomous driving

Zihao Liu, Xing Liu, Yizhai Zhang +1

Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varyi…

cs.CV2026

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

Yi Cui, Zilin Wang, Yijie Xu +5

Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only recognize common objects in cu…

physics.flu-dyn2026

Triggering of extreme events and coherent-structure modulation in wall-turbulence under cyclostationary forces

Ao Xu, Yun-Qian Bi, Heng-Dong Xi

Atmospheric gusts expose wall-bounded turbulence to severe unsteady forcing, triggering complex non-equilibrium dynamics and extreme aerodynamic loads. In this study, direct numeri…

cs.CL20262 cited

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

Haoxiang Sun, Tao Wang, Li Yuan +2

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of mo…