papers

Publications (16)

cs.CV2023

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Sihan Chen, Handong Li, Qunbo Wang +4

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient a…

cs.CV2023

COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

Sihan Chen, Xingjian He, Handong Li +3

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…

cs.CV2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

Handong Li, Zikang Liu, Longteng Guo +10

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception throug…

cs.CV2026

Thinking in Streaming Video

Zikang Liu, Longteng Guo, Handong Li +7

Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video re…

cs.CV2026

TimeThink: Reasoning with Time for Video LLMs

Handong Li, Longteng Guo, Zikang Liu +8

Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…

cs.CV2024

Explore the Limits of Omni-modal Pretraining at Scale

Yiyuan Zhang, Handong Li, Jing Liu +1

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining p…

q-fin.TR2021

Research on Portfolio Liquidation Strategy under Discrete Times

Qixuan Luo, Yu Shi, Handong Li

This paper presents an optimal strategy for portfolio liquidation under discrete time conditions. We assume that N risky assets held will be liquidated according to the same time i…

cond-mat.mtrl-sci2011

Current induced anisotropic magnetoresistance in topological insulator films

Jian Wang, Handong Li, Cui-Zu Chang +7

Topological insulators are insulating in the bulk but possess spin-momentum locked metallic surface states protected by time-reversal symmetry. The existence of these surface state…

cond-mat.mtrl-sci2014

Tunable Semimetallic State in Compressive-strained SrIrO3 Films Revealed by Transport Behaviors

Lunyong Zhang, Qifeng Liang, Ye Xiong +11

Orthorhombic SrIrO3 is a typical spin-orbit-coupling correlated metal that shows diversified physical properties under the external stimuli. Here nonlinear Hall effect and weakly t…

q-fin.RM2021

A Method for Predicting VaR by Aggregating Generalized Distributions Driven by the Dynamic Conditional Score

Shijia Song, Handong Li

Constructing a more effective value at risk (VaR) prediction model has long been a goal in financial risk management. In this paper, we propose a novel parametric approach and prov…

q-fin.RM2023

Unveiling Early Warning Signals of Systemic Risks in Banks: A Recurrence Network-Based Approach

Shijia Song, Handong Li

Bank crisis is challenging to define but can be manifested through bank contagion. This study presents a comprehensive framework grounded in nonlinear time series analysis to ident…

q-fin.RM2021

Value-at-Risk forecasting model based on normal inverse Gaussian distribution driven by dynamic conditional score

Shijia Song, Handong Li

Under the framework of dynamic conditional score, we propose a parametric forecasting model for Value-at-Risk based on the normal inverse Gaussian distribution (Hereinafter NIG-DCS…

cond-mat.mtrl-sci2011

Interplay between topological insulators and superconductors

Jian Wang, Cui-Zu Chang, Handong Li +8

Topological insulators are insulating in the bulk but possess metallic surface states protected by time-reversal symmetry. Here, we report a detailed electronic transport study in…

cs.CV2023

Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner

Zikang Liu, Sihan Chen, Longteng Guo +3

Large pre-trained multimodal models have demonstrated significant success in a range of downstream tasks, including image captioning, image-text retrieval, visual question answerin…

cond-mat.mtrl-sci2012

Growth and band alignment of Bi2Se3 topological insulator on H-terminated Si(111) van der Waals surface

Handong Li, Lei Gao, Hui Li +4

The van der Waals epitaxy of single crystalline Bi2Se3 film was achieved on hydrogen passivated Si(111) (H:Si) substrate by physical vapor deposition. Valence band structures of Bi…

cs.CV2025

Breaking the Encoder Barrier for Seamless Video-Language Understanding

Handong Li, Yiyuan Zhang, Longteng Guo +2

Most Video-Large Language Models (Video-LLMs) adopt an encoder-decoder framework, where a vision encoder extracts frame-wise features for processing by a language model. However, t…