From the 1 of 15 linked papers with an AI index.
15 papers
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
Rulin Zhou, Qiujie Song, Yujie Ma +10
Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a st…
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
Jiankun Wang, Yisen Gao, Ziwei Zhang +3
Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be p…
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Jiayi Luo, Hanxin Zhu, Chen Gao +5
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…
Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning
Haobo Zhang, Jiankun Wang, Suraj Rajendran +5
The paper introduces Dysco, a plug‑in technique for federated fine‑tuning of large models that dynamically allocates client‑specific LoRA subspaces to reduce interference caused by…
Structural Rationale Distillation via Reasoning Space Compression
Jialin Yang, Jiankun Wang, Jiajun Wu +3
When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure and strategy. Like a chef who…
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
Jiayi Luo, Jiayu Chen, Jiankun Wang +6
Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, motivating sparse attention techniques for impr…