Publications (76)
VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
Zhengfei Kuang, Rui Lin, Long Zhao +3
Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. I…
Epsilon-VAE: Denoising as Visual Decoding
Long Zhao, Sanghyun Woo, Ziyu Wan +6
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…
Simultaneous Polysomnography and Cardiotocography Reveal Temporal Correlation Between Maternal Obstructive Sleep Apnea and Fetal Hypoxia
Jingyu Wang, Donglin Xie, Jingying Ma +17
Background: Obstructive sleep apnea syndrome (OSAS) during pregnancy is common and can negatively affect fetal outcomes. However, studies on the immediate effects of maternal hypox…
A Two-Stage Allocation Scheme for Delay-Sensitive Services in Dense Vehicular Networks
Haojun Yang, Long Zhao, Lei Lei +1
Driven by the rapid development of wireless communication system, more and more vehicular services can be efficiently supported via vehicle-to-everything (V2X) communications. In o…
Generating Enhanced Negatives for Training Language-Based Object Detectors
Shiyu Zhao, Long Zhao, Vijay Kumar B. G +4
The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotatio…
Steering Prototypes with Prompt-tuning for Rehearsal-free Continual Learning
Zhuowei Li, Long Zhao, Zizhao Zhang +4
In the context of continual learning, prototypes-as representative class embeddings-offer advantages in memory conservation and the mitigation of catastrophic forgetting. However,…
View-Invariant, Occlusion-Robust Probabilistic Embedding for Human Pose
Ting Liu, Jennifer J. Sun, Long Zhao +6
Recognition of human poses and actions is crucial for autonomous systems to interact smoothly with people. However, cameras generally capture human poses in 2D as images and videos…
Semantic Graph Convolutional Networks for 3D Human Pose Regression
Long Zhao, Xi Peng, Yu Tian +2
In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convol…
Wormholes and the Thermodynamic Arrow of Time
Zhuo-Yu Xian, Long Zhao
In classical thermodynamics, heat cannot spontaneously pass from a colder system to a hotter system, which is called the thermodynamic arrow of time. However, if the initial states…
Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention
Yuxiao Chen, Long Zhao, Xi Peng +2
We propose a Dynamic Graph-Based Spatial-Temporal Attention (DG-STA) method for hand gesture recognition. The key idea is to first construct a fully-connected graph from a hand ske…
More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching
Yuxiao Chen, Jianbo Yuan, Long Zhao +4
Cross-modal attention mechanisms have been widely applied to the image-text matching task and have achieved remarkable improvements thanks to its capability of learning fine-graine…
Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding
Zizhao Zhang, Han Zhang, Long Zhao +3
Hierarchical structures are popular in recent vision transformers, however, they require sophisticated designs and massive datasets to work well. In this paper, we explore the idea…
Scenario Generation of Wind Farm Power for Real-Time System Operation
Trevor Werho, Junshan Zhang, Vijay Vittal +3
This work proposes a method of wind farm scenario generation to support real-time optimization tools and presents key findings therein. This work draws upon work from the literatur…
Unwinding Stochastic Order Flow: When to Warehouse Trades
Marcel Nutz, Kevin Webster, Long Zhao
We study how to unwind stochastic order flow with minimal transaction costs. Stochastic order flow arises, e.g., in the central risk book (CRB), a centralized trading desk that agg…
Bridging Saliency Detection to Weakly Supervised Object Detection Based on Self-paced Curriculum Learning
Dingwen Zhang, Deyu Meng, Long Zhao +1
Weakly-supervised object detection (WOD) is a challenging problems in computer vision. The key problem is to simultaneously infer the exact object locations in the training images…
Rethinking Kernel Methods for Node Representation Learning on Graphs
Yu Tian, Long Zhao, Xi Peng +1
Graph kernels are kernel methods measuring graph similarity and serve as a standard tool for graph classification. However, the use of kernel methods for node classification, which…
Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao, Xi Peng, Yu Tian +2
We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…
Detecting quantum chaos via pseudo-entropy
Song He, Pak Hang Chris Lau, Long Zhao
Quantum informatic quantities such as entanglement entropy are useful in detecting quantum phase transitions. Recently, a new entanglement measure called pseudo-entropy was propose…
Entanglement and Pseudo Entanglement Dynamics versus Fusion in CFT
Song He, Yu-Xuan Zhang, Long Zhao +1
The fusion rules and operator product expansion (OPE) serve as crucial tools in the study of operator algebras within conformal field theory (CFT). Building upon the vision of usin…
Distilling Vision-Language Models on Millions of Videos
Yue Zhao, Long Zhao, Xingyi Zhou +9
The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply…
Quantum chaos, scrambling and operator growth in deformed SYK models
Song He, Pak Hang Chris Lau, Zhuo-Yu Xian +1
In this work, we investigate the quantum chaos in various -deformed SYK models with finite , including the SYK, the supersymmetric SYK, and the SYK models.…
Timelike Entanglement Entropy in Higher Curvature Gravity
Zi-Xuan Zhao, Long Zhao, Song He
This work investigates holographic timelike entanglement entropy in higher curvature gravity, with a particular focus on Lovelock theories and on the role of excited states. For st…
InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
Yijie Zheng, Weijie Wu, Qingyun Li +7
Language-Guided object recognition in remote sensing imagery is crucial for large-scale mapping and automated data annotation. However, existing open-vocabulary and visual groundin…
COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality
Honglu Zhou, Asim Kadav, Aviv Shamsian +6
Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by m…
Energy spectrum of gravitational waves
Rong-Gen Cai, Xing-Yu Yang, Long Zhao
The energy spectrum of gravitational waves (GWs), which depicts the energy of GWs per unit volume of space per logarithmic frequency interval normalized to the critical density of…
The universality of islands outside the horizon
Song He, Yuan Sun, Long Zhao +1
We systematically calculate the quantum extremal surface (QES) associated with Hawking radiation for general -dimensional () asymptotically flat (or AdS) eternal black h…
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
Yuanhao Xiong, Long Zhao, Boqing Gong +5
Existing video-language pre-training methods primarily focus on instance-level alignment between video clips and captions via global contrastive learning but neglect rich fine-grai…
An Adaptive CSI Feedback Model Based on BiLSTM for Massive MIMO-OFDM Systems
Hongrui Shen, Long Zhao, Kan Zheng +2
Deep learning (DL)-based channel state information (CSI) feedback has the potential to improve the recovery accuracy and reduce the feedback overhead in massive multiple-input mult…
A Driving Intention Prediction Method Based on Hidden Markov Model for Autonomous Driving
Shiwen Liu, Kan Zheng, Long Zhao +1
In a mixed-traffic scenario where both autonomous vehicles and human-driving vehicles exist, a timely prediction of driving intentions of nearby human-driving vehicles is essential…
Twin-Timescale Radio Resource Management for Ultra-Reliable and Low-Latency Vehicular Networks
Haojun Yang, Kan Zheng, Long Zhao +1
To efficiently support safety-related vehicular applications, the ultra-reliable and low-latency communication (URLLC) concept has become an indispensable component of vehicular ne…
Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition
Qitong Wang, Long Zhao, Liangzhe Yuan +2
We are concerned with a challenging scenario in unpaired multiview video learning. In this case, the model aims to learn comprehensive multiview representations while the cross-vie…
Short-term Road Traffic Prediction based on Deep Cluster at Large-scale Networks
Lingyi Han, Kan Zheng, Long Zhao +2
Short-term road traffic prediction (STTP) is one of the most important modules in Intelligent Transportation Systems (ITS). However, network-level STTP still remains challenging du…
Enhanced Feature Extraction for IoT Network Intrusion Detection Using GNNs and KAN
Long Zhao, Shixun Ji, Bin Cheng +1
Recent advancements in the Internet of Things (IoT) emphasize the urgent need for advanced network security, as IoT networks feature dynamic topologies, imbalanced traffic, and com…
Deep Deformable Models: Learning 3D Shape Abstractions with Part Consistency
Di Liu, Long Zhao, Qilong Zhangli +3
The task of shape abstraction with semantic part consistency is challenging due to the complex geometries of natural objects. Recent methods learn to represent an object shape usin…
Island formula in Planck brane
Jing-Cheng Chang, Song He, Yu-Xiao Liu +1
Double holography offers a profound understanding of the island formula by describing a gravitational system on AdS coupled to a conformal field theory on ,…
Improved Transformer for High-Resolution GANs
Long Zhao, Zizhao Zhang, Ting Chen +2
Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making th…
Out-of-Domain Generalization from a Single Source: An Uncertainty Quantification Approach
Xi Peng, Fengchun Qiao, Long Zhao
We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…
Complexity/Action duality of shock wave geometry in a massive gravity theory
Yan-Gang Miao, Long Zhao
On the holographic complexity dual to the bulk action, we investigate the action growth for a shock wave geometry in a massive gravity theory within the Wheeler-De Witt (WDW) patch…
CR-GAN: Learning Complete Representations for Multi-view Generation
Yu Tian, Xi Peng, Long Zhao +2
Generating multi-view images from a single-view input is an essential yet challenging problem. It has broad applications in vision, graphics, and robotics. Our study indicates that…
Cooperative V2X for High Definition Map Transmission Based on Vehicle Mobility
Fangfei Wang, Dong Guan, Long Zhao +1
High-definition (HD) map transmission is considered as a key technology for automatic driving, which enables vehicles to obtain the precise road and surrounding environment informa…
Unified Visual Relationship Detection with Vision and Language Models
Long Zhao, Liangzhe Yuan, Boqing Gong +5
This work focuses on training a single visual relationship detector predicting over the union of label spaces from multiple datasets. Merging labels spanning different datasets cou…
Hierarchically Self-Supervised Transformer for Human Skeleton Representation Learning
Yuxiao Chen, Long Zhao, Jianbo Yuan +5
Despite the success of fully-supervised human skeleton sequence modeling, utilizing self-supervised pre-training for skeleton sequence representation learning has been an active fi…
Image Diffusion Preview with Consistency Solver
Fu-Yun Wang, Hao Zhou, Liangzhe Yuan +8
The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employin…
Are Multimodal Transformers Robust to Missing Modality?
Mengmeng Ma, Jian Ren, Long Zhao +2
Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly pr…
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
Long Zhao, Qinghe Wang, Jiaan Zhu +5
Reinforcement Learning from Human Feedback (RLHF) has become a key post-training paradigm for improving model quality. However, the synchronous three-stage RLHF pipeline is often b…
MSCENet: A Multi-Scale Correlation Enhanced Network for Anomaly Detection
Long Zhao, Shixun Ji, Zhipeng Wang +2
In the field of multivariate time series anomaly detection, against the backdrop of increasing data complexity and complex dependencies across multiple temporal scales, traditional…
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
Xiaoyu Zhu, Hao Zhou, Pengfei Xing +6
In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel…
Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
Long Zhao, Yuxiao Wang, Jiaping Zhao +7
We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…
VideoGLUE: Video General Understanding Evaluation of Foundation Models
Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14
We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…
EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images
Yijie Zheng, Weijie Wu, Bingyue Wu +4
While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these…
Holographic deformation of the entanglement entropy in (A)dS/CFT
Jing-Cheng Chang, Song He, Yu-Xiao Liu +1
In recent years, the holographic duality between -deformed conformal field theory (CFT) and Anti-de Sitter (AdS) spacetime with finite radial cutoff has received signific…
Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
Long Zhao, Ting Liu, Xi Peng +1
Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristic…
On the energy of gravitational waves
Rong-Gen Cai, Xing-Yu Yang, Long Zhao
The energy of gravitational waves is a fundamental problem in gravity theory. The existing descriptions for the energy of gravitational waves, such as the well-known Isaacson energ…
Comparative Assessment of Look-Ahead Economic Dispatch and Ramp Products for Grid Flexibility
Qian Zhang, Le Xie, Long Zhao +1
High renewable penetration increases the frequency and magnitude of net-load ramps, stressing real-time flexibility. Two commonly deployed remedies are look-ahead economic dispatch…
Video Creation by Demonstration
Yihong Sun, Hao Zhou, Liangzhe Yuan +7
We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physical…
Global Matching with Overlapping Attention for Optical Flow Estimation
Shiyu Zhao, Long Zhao, Zhixing Zhang +2
Optical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, t…
Beyond Lexical: A Semantic Retrieval Framework for Textual SearchEngine
Kuan Fang, Long Zhao, Zhan Shen +3
Search engine has become a fundamental component in various web and mobile applications. Retrieving relevant documents from the massive datasets is challenging for a search engine…
Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge
Long Zhao, Xi Peng, Yuxiao Chen +2
Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student)…
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking
Fangrui Zhu, Yunfeng Xi, Jianmo Ni +9
Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evalu…
VideoPrism: A Foundational Visual Encoder for Video Understanding
Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…
Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage
Wanshu Nie, Sujay V. Kumar, Junyu Chen +6
Recent advances in machine learning such as Long Short-Term Memory (LSTM) models and Transformers have been widely adopted in hydrological applications, demonstrating impressive pe…
SMIL: Multimodal Learning with Severely Missing Modality
Mengmeng Ma, Jian Ren, Long Zhao +3
A common assumption in multimodal learning is the completeness of training data, i.e., full modalities are available in all training examples. Although there exists research endeav…
Min-Max Latency Optimization Based on Sensed Position State Information in Internet of Vehicles
Pengzun Gao, Long Zhao, Kan Zheng +1
The dual-function radar communication (DFRC) is an essential technology in Internet of Vehicles (IoV). Consider that the road-side unit (RSU) employs the DFRC signals to sense the…
Limits of Semistatic Trading Strategies
Marcel Nutz, Johannes Wiesel, Long Zhao
We show that pointwise limits of semistatic trading strategies in discrete time are again semistatic strategies. The analysis is carried out in full generality for a two-period mod…
AuroraRL: Fast, Fault-Tolerant, and Cost-Efficient Reinforcement Learning over Decentralized Network
Chaoyi Ruan, Geng Luo, Xinyi Wan +12
LLM reinforcement learning (RL) requires frequent synchronization of large model parameters between the trainer and distributed rollout actors. High-throughput RL post-training the…
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
Zhuowei Li, Haizhou Shi, Yunhe Gao +7
Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded content…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Optimal Power Flow in Hybrid AC and Multi-terminal HVDC Networks with Offshore Wind Farm Integration Based on Semidefinite Programming
Yuhao Zhou, Long Zhao, Wei-Jen Lee +2
Multi-terminal high voltage direct current (MTHVDC) technology is a promising technology for the offshore wind farm integration, which requires the new control and operation scheme…
Cartoonish sketch-based face editing in videos using identity deformation transfer
Long Zhao, Fangda Han, Xi Peng +4
We address the problem of using hand-drawn sketches to create exaggerated deformations to faces in videos, such as enlarging the shape or modifying the position of eyes or mouth. T…
Learning to Learn Single Domain Generalization
Fengchun Qiao, Long Zhao, Xi Peng
We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…
Exploiting Unlabeled Data with Vision and Language Models for Object Detection
Shiyu Zhao, Zhixing Zhang, Samuel Schulter +5
Building robust and generic object detection frameworks requires scaling to larger label spaces and bigger training datasets. However, it is prohibitively costly to acquire annotat…
Reliable and Efficient Autonomous Driving: the Need for Heterogeneous Vehicular Networks
Kan Zheng, Qiang Zheng, Haojun Yang +3
Autonomous driving technology has been regarded as a promising solution to reduce road accidents and traffic congestion, as well as to optimize the usage of fuel and lane. Reliable…
Taming Self-Training for Open-Vocabulary Object Detection
Shiyu Zhao, Samuel Schulter, Long Zhao +5
Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pretrained vision and language models (VLMs). Howeve…
Martingale Schrödinger Bridges and Optimal Semistatic Portfolios
Marcel Nutz, Johannes Wiesel, Long Zhao
In a two-period financial market where a stock is traded dynamically and European options at maturity are traded statically, we study the so-called martingale Schrödinger bridge Q…
Box Re-Ranking: Unsupervised False Positive Suppression for Domain Adaptive Pedestrian Detection
Weijie Chen, Yilu Guo, Shicai Yang +7
False positive is one of the most serious problems brought by agnostic domain shift in domain adaptive pedestrian detection. However, it is impossible to label each box in countles…
TeleEval-OS: Performance evaluations of large language models for operations scheduling
Yanyan Wang, Yingying Wang, Junli Liang +10
The rapid advancement of large language models (LLMs) has significantly propelled progress in artificial intelligence, demonstrating substantial application potential across multip…