7 papers
SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
Ming Zhao, Wenhui Dong, Yang Zhang +23
Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets.…
How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
Yingjie He, Zhaolu Kang, Kehan Jiang +19
Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains underexplored. Sentence-level restor…
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
Yanxiang Huang, Guohua Gao, Zhaoyang Wei +1
Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the halluci…
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
Huaying Yuan, Jian Ni, Zheng Liu +7
Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of…
Temporal Consistency Constrained Transferable Adversarial Attacks with Background Mixup for Action Recognition
Ping Li, Jianan Ni, Bo Pang
Action recognition models using deep learning are vulnerable to adversarial examples, which are transferable across other models trained on the same data modality. Existing transfe…
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
Haode Qi, Cheng Qian, Jian Ni +6
In an enterprise Virtual Assistant (VA) system, intent classification is the crucial component that determines how a user input is handled based on what the user wants. The VA syst…