activity
20242026
most citedFrom Efficient Multimodal Models to World Models: A Survey

3 citations · 3 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CV2026

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

V Team, Wenyi Hong, Xiaotao Gu +94

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…

cs.CV2025

Component-aware Unsupervised Logical Anomaly Generation for Industrial Anomaly Detection

Xuan Tong, Yang Chang, Qing Zhao +9

Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits tra…

cs.CV2024

All rivers run into the sea: Unified Modality Brain-like Emotional Central Mechanism

Xinji Mai, Junxiong Lin, Haoran Wang +10

In the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emot…

cs.CV2024

Hi-EF: Benchmarking Emotion Forecasting in Human-interaction

Haoran Wang, Xinji Mai, Zeng Tao +6

Affective Forecasting is an psychology task that involves predicting an individual's future emotional responses, often hampered by reliance on external factors leading to inaccurac…

cs.LG20243 cited

From Efficient Multimodal Models to World Models: A Survey

Xinji Mai, Zeng Tao, Junxiong Lin +5

Multimodal Large Models (MLMs) are becoming a significant research focus, combining powerful large language models with multimodal learning to perform complex tasks across differen…

cs.CV2024

Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution

Junxiong Lin, Zeng Tao, Xuan Tong +10

The problem of blind image super-resolution aims to recover high-resolution (HR) images from low-resolution (LR) images with unknown degradation modes. Most existing methods model…