activity
20242026
collaborators

5 papers

cs.AI2026

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

Zhaoyang Zhang, Run Shao, Dongyue Wu +4

Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence leads, remains an open question…

cs.AI2026

M-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining

Rui Lv, Juncheng Mo, Tianyi Chu +11

Graphical User Interface (GUI) agent is pivotal to advancing intelligent human-computer interaction paradigms. Constructing powerful GUI agents necessitates the large-scale annotat…

cs.AI2025

SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing

Hongyi Jing, Jiafu Chen, Chen Rao +9

The existing Multimodal Large Language Models (MLLMs) for GUI perception have made great progress. However, the following challenges still exist in prior methods: 1) They model dis…

cs.CV2024

Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency

Yutong Wang, Jiajie Teng, Jiajiong Cao +4

As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-world online videos are often plagued by degradations such a…

cs.MM2024

MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding

Jiajie Teng, Huiyu Duan, Yucheng Zhu +2

Recent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which…