Publications (19)
CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model
Feiyang Wang, Xiaomin Yu, Wangyu Wu
Proving Rubik's Cube theorems at the high level represents a notable milestone in human-level spatial imagination and logic thinking and reasoning. Traditional Rubik's Cube robots,…
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
Chengwen Liu, Xiaomin Yu, Zhuoyue Chang +15
In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore ne…
Anisotropic Modality Align
Xiaomin Yu, Yijiang Li, Yuhui Zhang +8
Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of…
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
Xiaomin Yu, Yi Xin, Yuhui Zhang +12
Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of d…
Two-dimensional Noncommutative atom Gas with Anandan interaction
Xiaomin Yu, Kang Li
Landau like quantization of the Anandan system in a special electromagnetic field is studied. Unlike the cases of the AC system and the HMW system, the torques of the system on the…
Structured Evidence Selection for Weakly Supervised Video Anomaly Detection
Chenglizhao Chen, Tianxiang Nan, Wen Li +4
Weakly supervised video anomaly detection relies solely on video-level labels for training, making it difficult to accurately localize anomalous events in complex scenes. In real-w…
SpikEmo: Enhancing Emotion Recognition With Spiking Temporal Dynamics in Conversations
Xiaomin Yu, Feiyang Wang, Ziyue Qiao
In affective computing, the task of Emotion Recognition in Conversations (ERC) has emerged as a focal area of research. The primary objective of this task is to predict emotional s…
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
Tianyi Jiang, Arctanx An, Hengyi Feng +12
Human problem-solving is never the repetition of a single mindset, by which we mean a distinct mode of cognitive processing. When tackling a specific task, we do not rely on a sing…
QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining
Jun Han, Shuo Zhang, Wei Li +14
Financial markets are noisy and non-stationary, making alpha mining highly sensitive to backtest noise and regime shifts. While recent agentic frameworks improve automation, they o…
EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
Shuo Zhang, Chaofa Yuan, Ryan Guo +11
While LLM-based agents have shown promise for deep research, most existing approaches rely on fixed workflows that struggle to adapt to real-world, open-ended queries. Recent work…
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
Jianbo Lin, Xiaomin Yu, Yi Xin +7
Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on…
Similarity-Aware Multimodal Prompt Learning for Fake News Detection
Ye Jiang, Xiaomin Yu, Yimin Wang +3
The standard paradigm for fake news detection mainly utilizes text information to model the truthfulness of news. However, the discourse of online fake news is typically subtle and…
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5
Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, an…
Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities
Xiaomin Yu, Yezhaohui Wang, Yanfang Chen +5
In recent years, generative artificial intelligence models, represented by Large Language Models (LLMs) and Diffusion Models (DMs), have revolutionized content production methods.…
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks
Yang Liu, Xiaomin Yu, Gongyu Zhang +5
"A data scientist is tasked with developing a low-cost surgical VQA system for a 2-month workshop. Due to data sensitivity, she collects 50 hours of surgical video from a hospital,…
Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
Chenglizhao Chen, Yuchen Cao, Xinyu Liu +3
Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods genera…
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
Sudong Wang, Weiquan Huang, Xiaomin Yu +9
The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiab…
Text-Only Data Synthesis for Vision Language Model Training
Xiaomin Yu, Wenjie Zhang, Ziyue Qiao +2
Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is costly. In contrast, text data is…
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
Yang Liu, Ming Ma, Xiaomin Yu +5
Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…