activity
20242026
most citedOmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

1 citations · 1 across the 4 of their papers we have counts for

collaborators

10 papers

cs.CV2026

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

Panqi Yang, Haodong Jing, Jiahao Chao +5

Unified visual tokenization faces a fundamental trade-off between high-fidelity pixel reconstruction (spatial equivariance) and semantic abstraction (conceptual invariance). We att…

cs.AI2026

Breakthrough the Suboptimal Stable Point in Value-Factorization-Based Multi-Agent Reinforcement Learning

Lesong Tao, Yifei Wang, Haodong Jing +4

Value factorization, a popular paradigm in MARL, faces significant theoretical and algorithmic bottlenecks: its tendency to converge to suboptimal solutions remains poorly understo…

cs.CV2026

This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse

Junhao Jia, Jiaqi Wang, Yunyou Liu +4

Prototype networks provide an intrinsic case based explanation mechanism, but their interpretability is often undermined by prototype collapse, where multiple prototypes degenerate…

cs.AI20261 cited

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

Zhangquan Chen, Jiale Tao, Ruihuang Li +10

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still f…

cs.LG2026

We Need a More Robust Classifier: Dual Causal Learning Empowers Domain-Incremental Time Series Classification

Zhipeng Liu, Peibo Duan, Xuan Tang +6

The World Wide Web thrives on intelligent services that rely on accurate time series classification, which has recently witnessed significant progress driven by advances in deep le…

cs.CV2025

UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space

Panqi Yang, Haodong Jing, Nanning Zheng +1

In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehe…