2 citations · 2 across the 24 of their papers we have counts for
14 papers · 1 filter
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
Chen Ling, Hanqian Li, Dongnan Liu +9
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models t…
Vidu S1: A Real-Time Interactive Video Generation Model
Jintao Zhang, Kai Jiang, Jintao Chen +24
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment throug…
Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs
Yubo Gao, Haotian Wu, Hong Chen +8
Chain-of-Thought (CoT) has significantly enhanced LLM reasoning, yet often incurs substantial computational overhead due to "overthinking": generating excessively long rationales w…
Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding
Lina Zhang, Tonmoy Monsoor, Peizheng Li +23
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporal…
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
Jiale Liu, Jungang Li, Jieming Yu +13
Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a…
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization
Zhixin Lin, Jungang Li, Dongliang Xu +5
Mobile GUI agents powered by Multimodal Large Language Models (MLLMs) can execute complex tasks on mobile devices. Despite this progress, most existing systems still optimize task…