papers

Publications (76)

cs.CV2026

VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

Zhengfei Kuang, Rui Lin, Long Zhao +3

Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. I…

cs.CV2025

Epsilon-VAE: Denoising as Visual Decoding

Long Zhao, Sanghyun Woo, Ziyu Wan +6

In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…

eess.SP2025

Simultaneous Polysomnography and Cardiotocography Reveal Temporal Correlation Between Maternal Obstructive Sleep Apnea and Fetal Hypoxia

Jingyu Wang, Donglin Xie, Jingying Ma +17

Background: Obstructive sleep apnea syndrome (OSAS) during pregnancy is common and can negatively affect fetal outcomes. However, studies on the immediate effects of maternal hypox…

cs.IT2017

A Two-Stage Allocation Scheme for Delay-Sensitive Services in Dense Vehicular Networks

Haojun Yang, Long Zhao, Lei Lei +1

Driven by the rapid development of wireless communication system, more and more vehicular services can be efficiently supported via vehicle-to-everything (V2X) communications. In o…

cs.CV2024

Generating Enhanced Negatives for Training Language-Based Object Detectors

Shiyu Zhao, Long Zhao, Vijay Kumar B. G +4

The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotatio…

cs.LG2023

Steering Prototypes with Prompt-tuning for Rehearsal-free Continual Learning

Zhuowei Li, Long Zhao, Zizhao Zhang +4

In the context of continual learning, prototypes-as representative class embeddings-offer advantages in memory conservation and the mitigation of catastrophic forgetting. However,…

cs.CV2021

View-Invariant, Occlusion-Robust Probabilistic Embedding for Human Pose

Ting Liu, Jennifer J. Sun, Long Zhao +6

Recognition of human poses and actions is crucial for autonomous systems to interact smoothly with people. However, cameras generally capture human poses in 2D as images and videos…

cs.CV2020

Semantic Graph Convolutional Networks for 3D Human Pose Regression

Long Zhao, Xi Peng, Yu Tian +2

In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convol…

hep-th2020

Wormholes and the Thermodynamic Arrow of Time

Zhuo-Yu Xian, Long Zhao

In classical thermodynamics, heat cannot spontaneously pass from a colder system to a hotter system, which is called the thermodynamic arrow of time. However, if the initial states…

cs.CV2019

Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention

Yuxiao Chen, Long Zhao, Xi Peng +2

We propose a Dynamic Graph-Based Spatial-Temporal Attention (DG-STA) method for hand gesture recognition. The key idea is to first construct a fully-connected graph from a hand ske…

cs.CV2022

More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching

Yuxiao Chen, Jianbo Yuan, Long Zhao +4

Cross-modal attention mechanisms have been widely applied to the image-text matching task and have achieved remarkable improvements thanks to its capability of learning fine-graine…

cs.CV2021

Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding

Zizhao Zhang, Han Zhang, Long Zhao +3

Hierarchical structures are popular in recent vision transformers, however, they require sophisticated designs and massive datasets to work well. In this paper, we explore the idea…

stat.AP2021

Scenario Generation of Wind Farm Power for Real-Time System Operation

Trevor Werho, Junshan Zhang, Vijay Vittal +3

This work proposes a method of wind farm scenario generation to support real-time optimization tools and presents key findings therein. This work draws upon work from the literatur…

q-fin.TR2025

Unwinding Stochastic Order Flow: When to Warehouse Trades

Marcel Nutz, Kevin Webster, Long Zhao

We study how to unwind stochastic order flow with minimal transaction costs. Stochastic order flow arises, e.g., in the central risk book (CRB), a centralized trading desk that agg…

cs.CV2017

Bridging Saliency Detection to Weakly Supervised Object Detection Based on Self-paced Curriculum Learning

Dingwen Zhang, Deyu Meng, Long Zhao +1

Weakly-supervised object detection (WOD) is a challenging problems in computer vision. The key problem is to simultaneously infer the exact object locations in the training images…

cs.LG2019

Rethinking Kernel Methods for Node Representation Learning on Graphs

Yu Tian, Long Zhao, Xi Peng +1

Graph kernels are kernel methods measuring graph similarity and serve as a standard tool for graph classification. However, the use of kernel methods for node classification, which…

cs.CV2018

Learning to Forecast and Refine Residual Motion for Image-to-Video Generation

Long Zhao, Xi Peng, Yu Tian +2

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…

hep-th2025

Detecting quantum chaos via pseudo-entropy

Song He, Pak Hang Chris Lau, Long Zhao

Quantum informatic quantities such as entanglement entropy are useful in detecting quantum phase transitions. Recently, a new entanglement measure called pseudo-entropy was propose…

hep-th2024

Entanglement and Pseudo Entanglement Dynamics versus Fusion in CFT

Song He, Yu-Xuan Zhang, Long Zhao +1

The fusion rules and operator product expansion (OPE) serve as crucial tools in the study of operator algebras within conformal field theory (CFT). Building upon the vision of usin…

cs.CV2024

Distilling Vision-Language Models on Millions of Videos

Yue Zhao, Long Zhao, Xingyi Zhou +9

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply…

hep-th2022

Quantum chaos, scrambling and operator growth in deformed SYK models

Song He, Pak Hang Chris Lau, Zhuo-Yu Xian +1

In this work, we investigate the quantum chaos in various -deformed SYK models with finite , including the SYK, the supersymmetric SYK, and the SYK models.…

hep-th2025

Timelike Entanglement Entropy in Higher Curvature Gravity

Zi-Xuan Zhao, Long Zhao, Song He

This work investigates holographic timelike entanglement entropy in higher curvature gravity, with a particular focus on Lovelock theories and on the role of excited states. For st…

cs.CV2025

InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition

Yijie Zheng, Weijie Wu, Qingyun Li +7

Language-Guided object recognition in remote sensing imagery is crucial for large-scale mapping and automated data annotation. However, existing open-vocabulary and visual groundin…

cs.CV2022

COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality

Honglu Zhou, Asim Kadav, Aviv Shamsian +6

Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by m…

astro-ph.CO2021

Energy spectrum of gravitational waves

Rong-Gen Cai, Xing-Yu Yang, Long Zhao

The energy spectrum of gravitational waves (GWs), which depicts the energy of GWs per unit volume of space per logarithmic frequency interval normalized to the critical density of…

hep-th2022

The universality of islands outside the horizon

Song He, Yuan Sun, Long Zhao +1

We systematically calculate the quantum extremal surface (QES) associated with Hawking radiation for general -dimensional () asymptotically flat (or AdS) eternal black h…

cs.CV2024

Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding

Yuanhao Xiong, Long Zhao, Boqing Gong +5

Existing video-language pre-training methods primarily focus on instance-level alignment between video clips and captions via global contrastive learning but neglect rich fine-grai…

eess.SP2024

An Adaptive CSI Feedback Model Based on BiLSTM for Massive MIMO-OFDM Systems

Hongrui Shen, Long Zhao, Kan Zheng +2

Deep learning (DL)-based channel state information (CSI) feedback has the potential to improve the recovery accuracy and reduce the feedback overhead in massive multiple-input mult…

cs.LG2019

A Driving Intention Prediction Method Based on Hidden Markov Model for Autonomous Driving

Shiwen Liu, Kan Zheng, Long Zhao +1

In a mixed-traffic scenario where both autonomous vehicles and human-driving vehicles exist, a timely prediction of driving intentions of nearby human-driving vehicles is essential…

cs.IT2020

Twin-Timescale Radio Resource Management for Ultra-Reliable and Low-Latency Vehicular Networks

Haojun Yang, Kan Zheng, Long Zhao +1

To efficiently support safety-related vehicular applications, the ultra-reliable and low-latency communication (URLLC) concept has become an indispensable component of vehicular ne…

cs.CV2023

Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition

Qitong Wang, Long Zhao, Liangzhe Yuan +2

We are concerned with a challenging scenario in unpaired multiview video learning. In this case, the model aims to learn comprehensive multiview representations while the cross-vie…

cs.LG2019

Short-term Road Traffic Prediction based on Deep Cluster at Large-scale Networks

Lingyi Han, Kan Zheng, Long Zhao +2

Short-term road traffic prediction (STTP) is one of the most important modules in Intelligent Transportation Systems (ITS). However, network-level STTP still remains challenging du…

cs.CR2026

Enhanced Feature Extraction for IoT Network Intrusion Detection Using GNNs and KAN

Long Zhao, Shixun Ji, Bin Cheng +1

Recent advancements in the Internet of Things (IoT) emphasize the urgent need for advanced network security, as IoT networks feature dynamic topologies, imbalanced traffic, and com…

cs.CV2023

Deep Deformable Models: Learning 3D Shape Abstractions with Part Consistency

Di Liu, Long Zhao, Qilong Zhangli +3

The task of shape abstraction with semantic part consistency is challenging due to the complex geometries of natural objects. Recent methods learn to represent an object shape usin…

hep-th2024

Island formula in Planck brane

Jing-Cheng Chang, Song He, Yu-Xiao Liu +1

Double holography offers a profound understanding of the island formula by describing a gravitational system on AdS coupled to a conformal field theory on ,…

cs.CV2021

Improved Transformer for High-Resolution GANs

Long Zhao, Zizhao Zhang, Ting Chen +2

Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making th…

cs.CV2022

Out-of-Domain Generalization from a Single Source: An Uncertainty Quantification Approach

Xi Peng, Fengchun Qiao, Long Zhao

We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…

hep-th2017

Complexity/Action duality of shock wave geometry in a massive gravity theory

Yan-Gang Miao, Long Zhao

On the holographic complexity dual to the bulk action, we investigate the action growth for a shock wave geometry in a massive gravity theory within the Wheeler-De Witt (WDW) patch…

cs.CV2018

CR-GAN: Learning Complete Representations for Multi-view Generation

Yu Tian, Xi Peng, Long Zhao +2

Generating multi-view images from a single-view input is an essential yet challenging problem. It has broad applications in vision, graphics, and robotics. Our study indicates that…

eess.SP2019

Cooperative V2X for High Definition Map Transmission Based on Vehicle Mobility

Fangfei Wang, Dong Guan, Long Zhao +1

High-definition (HD) map transmission is considered as a key technology for automatic driving, which enables vehicles to obtain the precise road and surrounding environment informa…

cs.CV2023

Unified Visual Relationship Detection with Vision and Language Models

Long Zhao, Liangzhe Yuan, Boqing Gong +5

This work focuses on training a single visual relationship detector predicting over the union of label spaces from multiple datasets. Merging labels spanning different datasets cou…

cs.CV2023

Hierarchically Self-Supervised Transformer for Human Skeleton Representation Learning

Yuxiao Chen, Long Zhao, Jianbo Yuan +5

Despite the success of fully-supervised human skeleton sequence modeling, utilizing self-supervised pre-training for skeleton sequence representation learning has been an active fi…

cs.LG2026

Image Diffusion Preview with Consistency Solver

Fu-Yun Wang, Hao Zhou, Liangzhe Yuan +8

The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employin…

cs.CV2022

Are Multimodal Transformers Robust to Missing Modality?

Mengmeng Ma, Jian Ren, Long Zhao +2

Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly pr…

cs.AI2026

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

Long Zhao, Qinghe Wang, Jiaan Zhu +5

Reinforcement Learning from Human Feedback (RLHF) has become a key post-training paradigm for improving model quality. However, the synchronous three-stage RLHF pipeline is often b…

cs.CR2026

MSCENet: A Multi-Scale Correlation Enhanced Network for Anomaly Detection

Long Zhao, Shixun Ji, Zhipeng Wang +2

In the field of multivariate time series anomaly detection, against the backdrop of increasing data complexity and complex dependencies across multiple temporal scales, traditional…

cs.CV2024

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

Xiaoyu Zhu, Hao Zhou, Pengfei Xing +6

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel…

cs.CV2021

Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization

Long Zhao, Yuxiao Wang, Jiaping Zhao +7

We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…

cs.CV2024

VideoGLUE: Video General Understanding Evaluation of Foundation Models

Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14

We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…

cs.CV2026

EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images

Yijie Zheng, Weijie Wu, Bingyue Wu +4

While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these…

hep-th2025

Holographic deformation of the entanglement entropy in (A)dS/CFT

Jing-Cheng Chang, Song He, Yu-Xiao Liu +1

In recent years, the holographic duality between -deformed conformal field theory (CFT) and Anti-de Sitter (AdS) spacetime with finite radial cutoff has received signific…

cs.LG2020

Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness

Long Zhao, Ting Liu, Xi Peng +1

Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristic…

gr-qc2022

On the energy of gravitational waves

Rong-Gen Cai, Xing-Yu Yang, Long Zhao

The energy of gravitational waves is a fundamental problem in gravity theory. The existing descriptions for the energy of gravitational waves, such as the well-known Isaacson energ…

eess.SY2026

Comparative Assessment of Look-Ahead Economic Dispatch and Ramp Products for Grid Flexibility

Qian Zhang, Le Xie, Long Zhao +1

High renewable penetration increases the frequency and magnitude of net-load ramps, stressing real-time flexibility. Two commonly deployed remedies are look-ahead economic dispatch…

cs.CV2024

Video Creation by Demonstration

Yihong Sun, Hao Zhou, Liangzhe Yuan +7

We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physical…

cs.CV2022

Global Matching with Overlapping Attention for Optical Flow Estimation

Shiyu Zhao, Long Zhao, Zhixing Zhang +2

Optical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, t…

cs.IR2020

Beyond Lexical: A Semantic Retrieval Framework for Textual SearchEngine

Kuan Fang, Long Zhao, Zhan Shen +3

Search engine has become a fundamental component in various web and mobile applications. Retrieving relevant documents from the massive datasets is challenging for a search engine…

cs.CV2020

Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge

Long Zhao, Xi Peng, Yuxiao Chen +2

Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student)…

cs.CV2026

EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking

Fangrui Zhu, Yunfeng Xi, Jianmo Ni +9

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evalu…

cs.CV2025

VideoPrism: A Foundational Visual Encoder for Video Understanding

Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…

cs.LG2025

Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage

Wanshu Nie, Sujay V. Kumar, Junyu Chen +6

Recent advances in machine learning such as Long Short-Term Memory (LSTM) models and Transformers have been widely adopted in hydrological applications, demonstrating impressive pe…

cs.CV2021

SMIL: Multimodal Learning with Severely Missing Modality

Mengmeng Ma, Jian Ren, Long Zhao +3

A common assumption in multimodal learning is the completeness of training data, i.e., full modalities are available in all training examples. Although there exists research endeav…

eess.SP2022

Min-Max Latency Optimization Based on Sensed Position State Information in Internet of Vehicles

Pengzun Gao, Long Zhao, Kan Zheng +1

The dual-function radar communication (DFRC) is an essential technology in Internet of Vehicles (IoV). Consider that the road-side unit (RSU) employs the DFRC signals to sense the…

q-fin.MF2022

Limits of Semistatic Trading Strategies

Marcel Nutz, Johannes Wiesel, Long Zhao

We show that pointwise limits of semistatic trading strategies in discrete time are again semistatic strategies. The analysis is carried out in full generality for a two-period mod…

cs.DC2026

AuroraRL: Fast, Fault-Tolerant, and Cost-Efficient Reinforcement Learning over Decentralized Network

Chaoyi Ruan, Geng Luo, Xinyi Wan +12

LLM reinforcement learning (RL) requires frequent synchronization of large model parameters between the trainer and distributed rollout actors. High-throughput RL post-training the…

cs.CV2025

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering

Zhuowei Li, Haizhou Shi, Yunhe Gao +7

Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded content…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

eess.SY2019

Optimal Power Flow in Hybrid AC and Multi-terminal HVDC Networks with Offshore Wind Farm Integration Based on Semidefinite Programming

Yuhao Zhou, Long Zhao, Wei-Jen Lee +2

Multi-terminal high voltage direct current (MTHVDC) technology is a promising technology for the offshore wind farm integration, which requires the new control and operation scheme…

cs.CV2019

Cartoonish sketch-based face editing in videos using identity deformation transfer

Long Zhao, Fangda Han, Xi Peng +4

We address the problem of using hand-drawn sketches to create exaggerated deformations to faces in videos, such as enlarging the shape or modifying the position of eyes or mouth. T…

cs.CV2020

Learning to Learn Single Domain Generalization

Fengchun Qiao, Long Zhao, Xi Peng

We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…

cs.CV2022

Exploiting Unlabeled Data with Vision and Language Models for Object Detection

Shiyu Zhao, Zhixing Zhang, Samuel Schulter +5

Building robust and generic object detection frameworks requires scaling to larger label spaces and bigger training datasets. However, it is prohibitively costly to acquire annotat…

cs.NI2015

Reliable and Efficient Autonomous Driving: the Need for Heterogeneous Vehicular Networks

Kan Zheng, Qiang Zheng, Haojun Yang +3

Autonomous driving technology has been regarded as a promising solution to reduce road accidents and traffic congestion, as well as to optimize the usage of fuel and lane. Reliable…

cs.CV2024

Taming Self-Training for Open-Vocabulary Object Detection

Shiyu Zhao, Samuel Schulter, Long Zhao +5

Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pretrained vision and language models (VLMs). Howeve…

q-fin.MF2022

Martingale Schrödinger Bridges and Optimal Semistatic Portfolios

Marcel Nutz, Johannes Wiesel, Long Zhao

In a two-period financial market where a stock is traded dynamically and European options at maturity are traded statically, we study the so-called martingale Schrödinger bridge Q…

cs.CV2021

Box Re-Ranking: Unsupervised False Positive Suppression for Domain Adaptive Pedestrian Detection

Weijie Chen, Yilu Guo, Shicai Yang +7

False positive is one of the most serious problems brought by agnostic domain shift in domain adaptive pedestrian detection. However, it is impossible to label each box in countles…

cs.CL2025

TeleEval-OS: Performance evaluations of large language models for operations scheduling

Yanyan Wang, Yingying Wang, Junli Liang +10

The rapid advancement of large language models (LLMs) has significantly propelled progress in artificial intelligence, demonstrating substantial application potential across multip…