papers

Publications (26)

cs.CV2025

Frame-Voyager: Learning to Query Frames for Video Large Language Models

Sicheng Yu, Chengkai Jin, Huanyu Wang +9

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it…

cs.CV2026

Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs

Xuan Zhang, Cunxiao Du, Sicheng Yu +4

Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence length grows, posing challenges for the…

cs.CV2026

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

Yuan Zhang, Chun-Kai Fan, Sicheng Yu +6

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, perfor…

astro-ph.HE2022

Discovery of one neutron star candidate from radial velocity monitoring

Hailong Yuan, Song Wang, Zhongrui Bai +8

We report the discovery of one possible neutron star binary ( 0.8666 day) by using the LAMOST low-resolution spectroscopic data. The visible companion is a late A-ty…

cs.CV2025

An Uncertainty-aware DETR Enhancement Framework for Object Detection

Xingshu Chen, Sicheng Yu, Chong Cheng +2

This paper investigates the problem of object detection with a focus on improving both the localization accuracy of bounding boxes and explicitly modeling prediction uncertainty. C…

cs.CV2026

Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos

Gangjian Zhang, Jian Shu, Sicheng Yu +3

This paper addresses the challenge of reconstructing photorealistic and animatable 3D human avatars from monocular videos. While existing methods rely on combining per-subject opti…

cs.CV2025

3D Question Answering via only 2D Vision-Language Models

Fengyun Wang, Sicheng Yu, Jiawei Wu +3

Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, u…

cs.CV2025

VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction

Yu Hu, Chong Cheng, Sicheng Yu +2

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide a…

astro-ph.SR2022

ELM of ELM-WD: An extremely low mass hot star discovered in LAMOST survey

Hailong Yuan, Zhenwei Li, Zhongrui Bai +7

The Extremely Low Mass White Dwarfs (ELM WDs) and pre-ELM WDs are helium core white dwarfs with mass . Evolution simulations show that a lower mass limit for EL…

cs.CV2026

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

Kai Liu, Jungang Li, Yuchong Sun +13

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-…

cs.CL2021

NOAHQA: Numerical Reasoning with Interpretable Graph Question Answering Dataset

Qiyuan Zhang, Lei Wang, Sicheng Yu +4

While diverse question answering (QA) datasets have been proposed and contributed significantly to the development of deep learning models for QA tasks, the existing datasets fall…

cs.AI2025

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains

Wei Han, Geng Zhan, Sicheng Yu +2

O1/R1 style large reasoning models (LRMs) signal a substantial leap forward over conventional instruction-following LLMs. By applying test-time scaling to generate extended reasoni…

cs.CL2024

GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding

Cunxiao Du, Jing Jiang, Xu Yuanchen +8

Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs. In this study, we introduce GliDe and CaP…

cs.CV2026

Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes

Sicheng Yu, Dongxu Shen, Beizhen Zhao +2

Scaling monocular 3D Gaussian Splatting (3DGS) SLAM to kilometer-level outdoor environments poses two tightly coupled challenges: fragile long-term pose tracking and excessive memo…

cs.CV2026

Rotation-free Online Handwritten Character Recognition Using Linear Recurrent Units

Zhe Ling, Sicheng Yu, Danyu Yang

Online handwritten character recognition leverages stroke order and dynamic features, which generally provide higher accuracy and robustness compared with offline recognition. Howe…

cs.CV2024

OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation

Xiongwei Wu, Sicheng Yu, Ee-Peng Lim +1

In the realm of food computing, segmenting ingredients from images poses substantial challenges due to the large intra-class variance among the same ingredients, the emergence of n…

cs.CL2020

Context Modeling with Evidence Filter for Multiple Choice Question Answering

Sicheng Yu, Hao Zhang, Wei Jing +1

Multiple-Choice Question Answering (MCQA) is a challenging task in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that su…

cs.CV2025

RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes

Sicheng Yu, Chong Cheng, Yifan Zhou +2

3D Gaussian Splatting (3DGS) has become a popular solution in SLAM, as it can produce high-fidelity novel views. However, previous GS-based methods primarily target indoor scenes a…

cs.CV2025

Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping

Chong Cheng, Zijian Wang, Sicheng Yu +3

3D Gaussian Splatting (3DGS) has emerged as a core technique for 3D representation. Its effectiveness largely depends on precise camera poses and accurate point cloud initializatio…

cs.CV2025

Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps

Chong Cheng, Sicheng Yu, Zijian Wang +2

3D Gaussian Splatting (3DGS) has become a popular solution in SLAM due to its high-fidelity and real-time novel view synthesis performance. However, some previous 3DGS SLAM methods…

cs.CL2020

Counterfactual Variable Control for Robust and Interpretable Question Answering

Sicheng Yu, Yulei Niu, Shuohang Wang +2

Deep neural network based question answering (QA) models are neither robust nor explainable in many cases. For example, a multiple-choice QA model, tested without any input of ques…

cs.CV2026

Pseudo-View Enhancement via Confidence Fusion for Unposed Sparse-View Reconstruction

Beizhen Zhao, Sicheng Yu, Guanzhi Ding +2

3D scene reconstruction under unposed sparse viewpoints is a highly challenging yet practically important problem, especially in outdoor scenes due to complex lighting and scale va…

cs.GR2025

Wavelet-GS: 3D Gaussian Splatting with Wavelet Decomposition

Beizhen Zhao, Yifan Zhou, Sicheng Yu +2

3D Gaussian Splatting (3DGS) has revolutionized 3D scene reconstruction, which effectively balances rendering quality, efficiency, and speed. However, existing 3DGS approaches usua…

cs.CV2025

RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration

Chong Cheng, Yu Hu, Sicheng Yu +3

3D Gaussian Splatting (3DGS) has demonstrated its potential in reconstructing scenes from unposed images. However, optimization-based 3DGS methods struggle with sparse views due to…

cs.CV2026

MMGS: 10 Compressed 3DGS through Optimal Transport Aggregation based on Multi-view Ranking

Beizhen Zhao, Sicheng Yu, Ziran Yin +2

While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction, it suffers from significant overhead due to massive redundant primitives. Existing compression methods typi…

cs.CL2025

Reverse Modeling in Large Language Models

Sicheng Yu, Yuanchen Xu, Cunxiao Du +5

Humans are accustomed to reading and writing in a forward manner, and this natural bias extends to text understanding in auto-regressive large language models (LLMs). This paper in…