papers

Publications (19)

cs.RO2025

CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model

Feiyang Wang, Xiaomin Yu, Wangyu Wu

Proving Rubik's Cube theorems at the high level represents a notable milestone in human-level spatial imagination and logic thinking and reasoning. Traditional Rubik's Cube robots,…

cs.CV2026

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Chengwen Liu, Xiaomin Yu, Zhuoyue Chang +15

In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore ne…

cs.MM2026

Anisotropic Modality Align

Xiaomin Yu, Yijiang Li, Yuhui Zhang +8

Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of…

cs.CV2026

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

Xiaomin Yu, Yi Xin, Yuhui Zhang +12

Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of d…

physics.gen-ph2011

Two-dimensional Noncommutative atom Gas with Anandan interaction

Xiaomin Yu, Kang Li

Landau like quantization of the Anandan system in a special electromagnetic field is studied. Unlike the cases of the AC system and the HMW system, the torques of the system on the…

cs.CV2026

Structured Evidence Selection for Weakly Supervised Video Anomaly Detection

Chenglizhao Chen, Tianxiang Nan, Wen Li +4

Weakly supervised video anomaly detection relies solely on video-level labels for training, making it difficult to accurately localize anomalous events in complex scenes. In real-w…

cs.MM2024

SpikEmo: Enhancing Emotion Recognition With Spiking Temporal Dynamics in Conversations

Xiaomin Yu, Feiyang Wang, Ziyue Qiao

In affective computing, the task of Emotion Recognition in Conversations (ERC) has emerged as a focal area of research. The primary objective of this task is to predict emotional s…

cs.AI2026

Chain of Mindset: Reasoning with Adaptive Cognitive Modes

Tianyi Jiang, Arctanx An, Hengyi Feng +12

Human problem-solving is never the repetition of a single mindset, by which we mean a distinct mode of cognitive processing. When tackling a specific task, we do not rely on a sing…

q-fin.ST2026

QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining

Jun Han, Shuo Zhang, Wei Li +14

Financial markets are noisy and non-stationary, making alpha mining highly sensitive to backtest noise and regime shifts. While recent agentic frameworks improve automation, they o…

cs.AI2026

EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines

Shuo Zhang, Chaofa Yuan, Ryan Guo +11

While LLM-based agents have shown promise for deep research, most existing approaches rely on fixed workflows that struggle to adapt to real-world, open-ended queries. Recent work…

cs.AI2026

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

Jianbo Lin, Xiaomin Yu, Yi Xin +7

Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on…

cs.CL2023

Similarity-Aware Multimodal Prompt Learning for Fake News Detection

Ye Jiang, Xiaomin Yu, Yimin Wang +3

The standard paradigm for fake news detection mainly utilizes text information to model the truthfulness of news. However, the discourse of online fake news is typically subtle and…

cs.CV2026

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, an…

cs.CL2024

Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities

Xiaomin Yu, Yezhaohui Wang, Yanfang Chen +5

In recent years, generative artificial intelligence models, represented by Large Language Models (LLMs) and Diffusion Models (DMs), have revolutionized content production methods.…

cs.CV2024

ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks

Yang Liu, Xiaomin Yu, Gongyu Zhang +5

"A data scientist is tasked with developing a low-cost surgical VQA system for a 2-month workshop. Due to data sensitivity, she collects 50 hours of surgical video from a hospital,…

cs.CV2026

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

Chenglizhao Chen, Yuchen Cao, Xinyu Liu +3

Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods genera…

cs.CV2026

Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

Sudong Wang, Weiquan Huang, Xiaomin Yu +9

The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiab…

cs.AI2026

Text-Only Data Synthesis for Vision Language Model Training

Xiaomin Yu, Wenjie Zhang, Ziyue Qiao +2

Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is costly. In contrast, text data is…

cs.CV2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu +5

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…