activity
20242026
collaborators

7 papers

cs.CV2026

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis

Shihao Yuan, Yuanze Li, Ruyi Zhang +2

Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue t…

cs.LG2026

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

Liu Hanqing, Jianjun Cao, Yuanze Li +1

Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this to…

cs.CV2025

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

Zichen Tang, Haihong E, Rongjin Li +18

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to exis…

cs.CV2025

Triad: Empowering LMM-based Anomaly Detection with Vision Expert-guided Visual Tokenizer and Manufacturing Process

Yuanze Li, Shihao Yuan, Haolin Wang +5

Although recent methods have tried to introduce large multimodal models (LMMs) into industrial anomaly detection (IAD), their generalization in the IAD field is far inferior to tha…

cs.CV2025

Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection

Yuanze Li, Haolin Wang, Shihao Yuan +6

Due to the training configuration, traditional industrial anomaly detection (IAD) methods have to train a specific model for each deployment scenario, which is insufficient to meet…

cs.CV2024

Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective

Yuanze Li, Chun-Mei Feng, Qilong Wang +2

Human beings can leverage knowledge from relative tasks to improve learning on a primary task. Similarly, multi-task learning methods suggest using auxiliary tasks to enhance a neu…