3 papers
cs.CV2026
Beyond Single-Sample: Reliable Multi-Sample Distillation for Video Understanding
Songlin Li, Xin Zhu, Zechao Guan +2
Traditional black-box distillation for Large Vision-Language Models (LVLMs) typically relies on a single teacher response per input, which often yields high-variance responses and…
cs.RO2025
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
Feng Yan, Fanfan Liu, Liming Zheng +5
Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial i…
cs.CV2025
TopoDiT-3D: Topology-Aware Diffusion Transformer with Bottleneck Structure for 3D Point Cloud Generation
Zechao Guan, Feng Yan, Shuai Du +2
Recent advancements in Diffusion Transformer (DiT) models have significantly improved 3D point cloud generation. However, existing methods primarily focus on local feature extracti…