4 papers
Allocation Before Ranking: Decoupled Token Compression for OmniLLMs
Zhenghui Guo, Yilin Yang, Yuanbin Man +5
Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The…
Splaxel: Efficient Distributed Training of 3D Gaussian Splatting for Large-scale Scene Reconstruction via Pixel-level Communication
Wenqi Jia, Zhewen Hu, Ying Huang +10
3D Gaussian Splatting (3DGS) enables high-fidelity and real-time 3D scene reconstruction, but scaling training to large-scale scenes requires optimizing hundreds of millions of Gau…
HII-DPO: Eliminate Hallucination via Accurate Hallucination-Inducing Counterfactual Images
Yilin Yang, Zhenghui Guo, Yuke Wang +3
Large Vision-Language Models (VLMs) have achieved remarkable success across diverse multimodal tasks but remain vulnerable to hallucinations rooted in inherent language bias. Despi…
SDiT: Semantic Region-Adaptive for Diffusion Transformers
Bowen Lin, Fanjiang Ye, Yihua Liu +7
Diffusion Transformers (DiTs) achieve state-of-the-art performance in text-to-image synthesis but remain computationally expensive due to the iterative nature of denoising and the…