Publications (26)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
Sicheng Yu, Chengkai Jin, Huanyu Wang +9
Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it…
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
Xuan Zhang, Cunxiao Du, Sicheng Yu +4
Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence length grows, posing challenges for the…
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
Yuan Zhang, Chun-Kai Fan, Sicheng Yu +6
Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, perfor…
Discovery of one neutron star candidate from radial velocity monitoring
Hailong Yuan, Song Wang, Zhongrui Bai +8
We report the discovery of one possible neutron star binary ( 0.8666 day) by using the LAMOST low-resolution spectroscopic data. The visible companion is a late A-ty…
An Uncertainty-aware DETR Enhancement Framework for Object Detection
Xingshu Chen, Sicheng Yu, Chong Cheng +2
This paper investigates the problem of object detection with a focus on improving both the localization accuracy of bounding boxes and explicitly modeling prediction uncertainty. C…
Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos
Gangjian Zhang, Jian Shu, Sicheng Yu +3
This paper addresses the challenge of reconstructing photorealistic and animatable 3D human avatars from monocular videos. While existing methods rely on combining per-subject opti…
3D Question Answering via only 2D Vision-Language Models
Fengyun Wang, Sicheng Yu, Jiawei Wu +3
Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, u…
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
Yu Hu, Chong Cheng, Sicheng Yu +2
Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide a…
ELM of ELM-WD: An extremely low mass hot star discovered in LAMOST survey
Hailong Yuan, Zhenwei Li, Zhongrui Bai +7
The Extremely Low Mass White Dwarfs (ELM WDs) and pre-ELM WDs are helium core white dwarfs with mass . Evolution simulations show that a lower mass limit for EL…
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
Kai Liu, Jungang Li, Yuchong Sun +13
This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-…
NOAHQA: Numerical Reasoning with Interpretable Graph Question Answering Dataset
Qiyuan Zhang, Lei Wang, Sicheng Yu +4
While diverse question answering (QA) datasets have been proposed and contributed significantly to the development of deep learning models for QA tasks, the existing datasets fall…
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
Wei Han, Geng Zhan, Sicheng Yu +2
O1/R1 style large reasoning models (LRMs) signal a substantial leap forward over conventional instruction-following LLMs. By applying test-time scaling to generate extended reasoni…
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
Cunxiao Du, Jing Jiang, Xu Yuanchen +8
Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs. In this study, we introduce GliDe and CaP…
Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes
Sicheng Yu, Dongxu Shen, Beizhen Zhao +2
Scaling monocular 3D Gaussian Splatting (3DGS) SLAM to kilometer-level outdoor environments poses two tightly coupled challenges: fragile long-term pose tracking and excessive memo…
Rotation-free Online Handwritten Character Recognition Using Linear Recurrent Units
Zhe Ling, Sicheng Yu, Danyu Yang
Online handwritten character recognition leverages stroke order and dynamic features, which generally provide higher accuracy and robustness compared with offline recognition. Howe…
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
Xiongwei Wu, Sicheng Yu, Ee-Peng Lim +1
In the realm of food computing, segmenting ingredients from images poses substantial challenges due to the large intra-class variance among the same ingredients, the emergence of n…
Context Modeling with Evidence Filter for Multiple Choice Question Answering
Sicheng Yu, Hao Zhang, Wei Jing +1
Multiple-Choice Question Answering (MCQA) is a challenging task in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that su…
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
Sicheng Yu, Chong Cheng, Yifan Zhou +2
3D Gaussian Splatting (3DGS) has become a popular solution in SLAM, as it can produce high-fidelity novel views. However, previous GS-based methods primarily target indoor scenes a…
Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping
Chong Cheng, Zijian Wang, Sicheng Yu +3
3D Gaussian Splatting (3DGS) has emerged as a core technique for 3D representation. Its effectiveness largely depends on precise camera poses and accurate point cloud initializatio…
Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
Chong Cheng, Sicheng Yu, Zijian Wang +2
3D Gaussian Splatting (3DGS) has become a popular solution in SLAM due to its high-fidelity and real-time novel view synthesis performance. However, some previous 3DGS SLAM methods…
Counterfactual Variable Control for Robust and Interpretable Question Answering
Sicheng Yu, Yulei Niu, Shuohang Wang +2
Deep neural network based question answering (QA) models are neither robust nor explainable in many cases. For example, a multiple-choice QA model, tested without any input of ques…
Pseudo-View Enhancement via Confidence Fusion for Unposed Sparse-View Reconstruction
Beizhen Zhao, Sicheng Yu, Guanzhi Ding +2
3D scene reconstruction under unposed sparse viewpoints is a highly challenging yet practically important problem, especially in outdoor scenes due to complex lighting and scale va…
Wavelet-GS: 3D Gaussian Splatting with Wavelet Decomposition
Beizhen Zhao, Yifan Zhou, Sicheng Yu +2
3D Gaussian Splatting (3DGS) has revolutionized 3D scene reconstruction, which effectively balances rendering quality, efficiency, and speed. However, existing 3DGS approaches usua…
RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration
Chong Cheng, Yu Hu, Sicheng Yu +3
3D Gaussian Splatting (3DGS) has demonstrated its potential in reconstructing scenes from unposed images. However, optimization-based 3DGS methods struggle with sparse views due to…
MMGS: 10 Compressed 3DGS through Optimal Transport Aggregation based on Multi-view Ranking
Beizhen Zhao, Sicheng Yu, Ziran Yin +2
While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction, it suffers from significant overhead due to massive redundant primitives. Existing compression methods typi…
Reverse Modeling in Large Language Models
Sicheng Yu, Yuanchen Xu, Cunxiao Du +5
Humans are accustomed to reading and writing in a forward manner, and this natural bias extends to text understanding in auto-regressive large language models (LLMs). This paper in…