26 papers
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
Chen Ling, Hanqian Li, Dongnan Liu +9
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models t…
Vidu S1: A Real-Time Interactive Video Generation Model
Jintao Zhang, Kai Jiang, Jintao Chen +24
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment throug…
Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs
Yubo Gao, Haotian Wu, Hong Chen +8
Chain-of-Thought (CoT) has significantly enhanced LLM reasoning, yet often incurs substantial computational overhead due to "overthinking": generating excessively long rationales w…
SpikingMoE: SDPrompt-Guided Dynamic Expert Fusion in Spiking Neural Networks
Yukai Yang, Chenxi Qin, Jungang Li +3
Spiking Neural Networks (SNNs) provide an energy-efficient paradigm for visual recognition. We present SpikingMoE, which integrates a spike-driven Transformer with a Mixture-of-Exp…
Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding
Lina Zhang, Tonmoy Monsoor, Peizheng Li +23
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporal…
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
Jiale Liu, Jungang Li, Jieming Yu +13
Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a…