1 citations · 1 across the 7 of their papers we have counts for
7 papers
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
Jiazheng Xing, Hangjie Yuan, Lingling Cai +9
Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified t…
COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
Darryl Cherian Jacob, Xinyu Liu, Kai Wang +1
Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suff…
AIvaluateXR: An Evaluation Framework for on-Device AI in XR with Benchmarking Results
Dawar Khan, Xinyu Liu, Omar Mena +3
The deployment of large language models (LLMs) on extended reality (XR) devices has great potential to advance the field of human-AI interaction. In the case of direct, on-device m…
ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
Dawar Khan, Alexandre Kouyoumdjian, Xinyu Liu +3
We present ClickAIXR, a novel on-device framework for multimodal vision-language interaction with objects in extended reality (XR). Unlike prior systems that rely on cloud-based AI…
XAttnRes: Cross-Stage Attention Residuals for Medical Image Segmentation
Xinyu Liu, Qing Xu, Zhen Chen
In the field of Large Language Models (LLMs), Attention Residuals have recently demonstrated that learned, selective aggregation over all preceding layer outputs can outperform fix…