papers

Publications (135)

eess.SP2025

UAV-Enabled Joint Sensing, Communication, Powering and Backhaul Transmission in Maritime Monitoring Networks

Bohan Li, Jiahao Liu, Yujun Liang +6

This paper addresses the challenge of energy-constrained maritime monitoring networks by proposing an unmanned aerial vehicle (UAV)-enabled integrated sensing, communication, power…

cs.IT2022

Multicarrier-Division Duplex for Solving the Channel Aging Problem in Massive MIMO Systems

Bohan Li, Lie-Liang Yang, Robert G Maunder +2

The separation of training and data transmission as well as the frequent uplink/downlink (UL/DL) switching make time-division duplex (TDD)-based massive multiple-input multiple-out…

cs.CV2026

OmniNWM: Omniscient Driving Navigation World Models

Bohan Li, Zhuang Ma, Dalong Du +10

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods are typically restricted to frag…

cs.CV2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

Wenyao Zhang, Hongsi Liu, Bohan Li +7

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this cha…

cs.CL2024

Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits

Bohan Li, Jiannan Guan, Longxu Dou +12

The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality de…

cs.CV2025

Light-X: Generative 4D Video Rendering with Camera and Illumination Control

Tianqi Liu, Zhaoxi Chen, Zihao Huang +8

Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighti…

physics.optics2019

All Optical Neural Network with Nonlinear Activation Functions

Ying Zuo, Bohan Li, Yujun Zhao +6

Artificial neural networks (ANNs) have now been widely used for industry applications and also played more important roles in fundamental researches. Although most ANN hardware sys…

cs.CL2024

A Two-Stage Framework with Self-Supervised Distillation For Cross-Domain Text Classification

Yunlong Feng, Bohan Li, Libo Qin +2

Cross-domain text classification aims to adapt models to a target domain that lacks labeled data. It leverages or reuses rich labeled data from the different but related source dom…

cs.CL2023

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

Yihan Wu, Junliang Guo, Xu Tan +7

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of…

cs.CV2023

ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

Wenwen Yu, Chengquan Zhang, Haoyu Cao +24

Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…

math.OC2025

Optimal investment problem in a renewal risk model with generalized Erlang distributed interarrival times

Linlin Tian, Yixuan Tian, Bohan Li +1

This paper explores the optimal investment problem of a renewal risk model with generalized Erlang distributed interarrival times. The phases of the Erlang interarrival time is ass…

cs.CL2023

MixPro: Simple yet Effective Data Augmentation for Prompt-based Learning

Bohan Li, Longxu Dou, Yutai Hou +5

Prompt-based learning has shown considerable promise in reformulating various downstream tasks as cloze problems by combining original input with a predetermined template. This app…

math.OC2023

Optimal Monotone Mean-Variance Problem in a Catastrophe Insurance Model

Bohan Li, Junyi Guo, Xiaoqing Liang

This paper explores an optimal investment and reinsurance problem involving both ordinary and catastrophe insurance businesses. The catastrophic events are modeled as following a c…

cs.SD2025

ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models

Bohan Li, Wenbin Huang, Yuhang Qiu +7

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted int…

physics.optics2022

Correlated self-heterodyne method for ultra-low-noise laser linewidth measurements

Zhiquan Yuan, Heming Wang, Peng Liu +9

Narrow-linewidth lasers are important to many applications spanning precision metrology to sensing systems. Characterization of these lasers requires precise measurements of their…

cs.CL2023

MetaPrompting: Learning to Learn Better Prompts

Yutai Hou, Hongyuan Dong, Xinghao Wang +2

Prompting method is regarded as one of the crucial progress for few-shot nature language processing. Recent research on prompting moves from discrete tokens based ``hard prompts''…

cs.CV2026

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

Xueqing Yu, Bohan Li, Yan Li +1

Recent Vision-Language Models (VLMs) have made remarkable progress in multimodal understanding tasks, yet their evaluation on long video understanding remains unreliable. Due to li…

cs.HC2024

Connecting Dreams with Visual Brainstorming Instruction

Yasheng Sun, Bohan Li, Mingchen Zhuge +4

Recent breakthroughs in understanding the human brain have revealed its impressive ability to efficiently process and interpret human thoughts, opening up possibilities for interve…

eess.AS2025

Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding

Bohan Li, Hankun Wang, Situo Zhang +2

The auto-regressive architecture, like GPTs, is widely used in modern Text-to-Speech (TTS) systems. However, it incurs substantial inference time, particularly due to the challenge…

cs.SD2025

Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency

Haoran Wang, Guanyu Chen, Bohan Li +5

Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexpl…

cs.SE2025

ReF Decompile: Relabeling and Function Call Enhanced Decompile

Yunlong Feng, Bohan Li, Xiaoming Shi +2

The goal of decompilation is to convert compiled low-level code (e.g., assembly code) back into high-level programming languages, enabling analysis in scenarios where source code i…

cs.DC2026

FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs

Bohan Li, Shengmin Li, Xinyu Shi +3

Graph Convolutional Networks (GCNs) are widely adopted for tasks involving relational or graph-structured data and can be formulated as two-stage sparse-dense matrix multiplication…

cs.CV2025

TAPTRv2: Attention-based Position Update Improves Tracking Any Point

Hongyang Li, Hao Zhang, Shilong Liu +5

In this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DE…

math.OC2023

Linear Quadratic Extended Mean Field Games and Control Problems

Alain Bensoussan, Bohan Li, Sheung Chi Phillip Yam

We provide a thorough study of a general class of linear-quadratic extended mean field games and control problems in any dimensions where the mean field terms are allowed to be unb…

cs.SD2021

DelightfulTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2021

Yanqing Liu, Zhihang Xu, Gang Wang +6

This paper describes the Microsoft end-to-end neural text to speech (TTS) system: DelightfulTTS for Blizzard Challenge 2021. The goal of this challenge is to synthesize natural and…

cs.CV2025

Closed-Loop Unsupervised Representation Disentanglement with -VAE Distillation and Diffusion Probabilistic Feedback

Xin Jin, Bohan Li, BAAO Xie +5

Representation disentanglement may help AI fundamentally understand the real world and thus benefit both discrimination and generation tasks. It currently has at least three unreso…

eess.SP2025

Joint Beamforming and Compressed Sensing for Uplink Grant-Free Access

Guoqing Xia, Pei Xiao, Bohan Li +2

Compressed sensing (CS)-based techniques have been widely applied in the grant-free non-orthogonal multiple access (NOMA) to a single-antenna base station (BS). In this paper, we c…

cs.SD2021

AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style

Yuzi Yan, Xu Tan, Bohan Li +6

While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e…

eess.IV2023

Multi-rate adaptive transform coding for video compression

Lyndon R. Duong, Bohan Li, Cheng Chen +1

Contemporary lossy image and video coding standards rely on transform coding, the process through which pixels are mapped to an alternative representation to facilitate efficient d…

cond-mat.mtrl-sci2026

kALDo 2.0: Scalable Thermal Transport from First Principles and Machine Learning Potentials

Giuseppe Barbalinardo, Zekun Chen, Dylan Folkner +5

We introduce kALDo2.0, an open-source Python package for computing vibrational, elastic, and thermal transport properties of solids from first principles and machine-learned intera…

hep-th2023

On low rank 4d SCFTs

Bohan Li, Dan Xie, Wenbin Yan

There are two major ways of constructing 4d superconformal field theories (SCFTs): the first one is putting a 6d theory on a punctured Riemann surface (clas…

physics.chem-ph2026

Harnessing AtomisticSkills for Agentic Atomistic Research

Bowen Deng, Bohan Li, Matthew Cox +20

Computational materials science and chemistry span vast knowledge domains and fractured software ecosystems. Although large language models (LLMs) have demonstrated research capabi…

cs.CV2026

GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

GigaBrain Team, Boyuan Wang, Bohan Li +23

Vision-language-action (VLA) models that directly predict multi-step action chunks from current observations face inherent limitations due to constrained scene understanding and we…

eess.AS2025

Recent Advances in Discrete Speech Tokens: A Review

Yiwei Guo, Zhihan Li, Hankun Wang +7

The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech repr…

eess.AS2021

AdaSpeech: Adaptive Text to Speech for Custom Voice

Mingjian Chen, Xu Tan, Bohan Li +4

Custom voice, a specific text to speech (TTS) service in commercial speech platforms, aims to adapt a source TTS model to synthesize personal voice for a target speaker using few s…

cs.SD2024

On the Effectiveness of Acoustic BPE in Decoder-Only TTS

Bohan Li, Feiyu Shen, Yiwei Guo +3

Discretizing speech into tokens and generating them by a decoder-only model have been a promising direction for text-to-speech (TTS) and spoken language modeling (SLM). To shorten…

physics.optics2023

Engineered zero-dispersion microcombs using CMOS-ready photonics

Qing-Xin Ji, Warren Jin, Lue Wu +14

Normal group velocity dispersion (GVD) microcombs offer high comb line power and high pumping efficiency compared to bright pulse microcombs. The recent demonstration of normal GVD…

eess.IV2021

A Technical Overview of AV1

Jingning Han, Bohan Li, Debargha Mukherjee +12

The AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than 30% reduction in bit-rate compared to its predecessor VP9 for the sam…

eess.AS2026

CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate

Hankun Wang, Yiwei Guo, Chongtian Shao +2

Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number o…

cs.SC2023

Efficient Local Search for Nonlinear Real Arithmetic

Zhonghan Wang, Bohua Zhan, Bohan Li +1

Local search has recently been applied to SMT problems over various arithmetic theories. Among these, nonlinear real arithmetic poses special challenges due to its uncountable solu…

cs.LG2021

Follow Your Path: a Progressive Method for Knowledge Distillation

Wenxian Shi, Yuxuan Song, Hao Zhou +2

Deep neural networks often have a huge number of parameters, which posts challenges in deployment in application scenarios with limited memory and computation capacity. Knowledge d…

cs.LG2019

A Surprisingly Effective Fix for Deep Latent Variable Modeling of Text

Bohan Li, Junxian He, Graham Neubig +2

When trained effectively, the Variational Autoencoder (VAE) is both a powerful language model and an effective representation learning framework. In practice, however, VAEs are tra…

cs.SD2025

Next Tokens Denoising for Speech Synthesis

Yanqing Liu, Ruiqing Xue, Chong Zhang +7

While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, c…

cs.LO2024

SMT-Layout: A MaxSMT-based Approach Supporting Real-time Interaction of Real-world GUI Layout

Bohan Li, Dawei Li, Ming Fu +1

Leveraging the flexible expressive ability of (Max)SMT and the powerful solving ability of SMT solvers, we propose a novel layout model named SMT-Layout. SMT-Layout is the first co…

physics.optics2022

Self-injection-locked second-harmonic integrated source

Jingwei Ling, Jeremy Staffa, Heming Wang +12

High coherence visible and near-visible laser sources are centrally important to the operation of advanced position/navigation/timing systems as well as classical/quantum sensing s…

cs.DB2025

ODIN: Object Density Aware Index for CkNN Queries over Moving Objects on Road Networks

Ziqiang Yu, Xiaohui Yu, Tao Zhou +3

We study the problem of processing continuous k nearest neighbor (CkNN) queries over moving objects on road networks, which is an essential operation in a variety of applications.…

cs.CV2026

Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

Songtao Tian, Guhan Chen, Bohan Li +2

Consistency distillation has significantly accelerated diffusion-model inference, but its sampling dynamics remain underexplored. We reveal an asymmetry: although Logit-Normal samp…

cond-mat.mes-hall2024

Unconventional superconductivity in magic-strain graphene superlattices

Qingxiang Ji, Bohan Li, Johan Christensen +2

Extensive investigations on the Moiré magic-angle have been conducted in twisted bilayer graphene, unlocking the mystery of unconventional superconductivity and insulating states.…

cs.CV2024

Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification

Bohan Li, Xiao Xu, Xinghao Wang +6

Existing image augmentation methods consist of two categories: perturbation-based methods and generative methods. Perturbation-based methods apply pre-defined perturbations to augm…

cs.SD2026

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Bohan Li, Shi Lian, Hankun Wang +6

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…

cs.LG2026

Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets

Zixiong Yu, Guhan Chen, Jianfa Lai +2

Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free architectures. While prior work ha…

cs.CV2026

Light of Normals: Unified Feature Representation for Universal Photometric Stereo

Houyuan Chen, Hong Li, Chongjie Ye +11

Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination model…

quant-ph2023

On the equivalence between squeezing and entanglement potential for two-mode Gaussian states

Bohan Li, Aritra Das, Spyros Tserkis +3

The maximum amount of entanglement achievable under passive transformations by continuous-variable states is called the entanglement potential. Recent work has demonstrated that th…

cs.CV2025

DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation

Jiazhe Guo, Yikang Ding, Xiwu Chen +8

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-sce…

cs.CV2026

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

Bohan Li, Shuojue Yang, Baorui Peng +10

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must prec…

math.OC2022

Optimal investment and reinsurance policies for the Cram{é}r-Lundberg risk model under monotone mean-variance preference

Bohan Li, Junyi Guo, Linlin Tian

In this paper, an optimization problem for the monotone mean-variance(MMV) criterion is considered in the perspective of the insurance company. The MMV criterion is an amended vers…

cs.SD2026

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

Ziyang Ma, Ruiyang Xu, Yinghao Ma +9

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Cha…

q-fin.RM2025

Mean Field Analysis of Mutual Insurance Market

Bohan Li, Wenyuan Li, Kenneth Tsz Hin Ng +1

A mutual insurance company (MIC) is a type of consumer cooperative owned by its policyholders. By purchasing insurance from an MIC, policyholders effectively become member-owners o…

cs.CL2025

U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-In-A-Haystack

Yunfan Gao, Yun Xiong, Wenlong Wu +3

Recent advancements in Large Language Models (LLMs) have expanded their context windows to unprecedented lengths, sparking debates about the necessity of Retrieval-Augmented Genera…

eess.SP2022

Heterogeneous graph neural network for power allocation in multicarrier-division duplex cell-free massive MIMO systems

Bohan Li, Lie-Liang Yang, Robert G Maunder +2

In-band full duplex cell-free (CF) systems suffer from severe self-interference and cross-link interference, especially when CF systems are operated in distributed way. To this end…

cs.CV2025

OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

Bohan Li, Xin Jin, Jianan Wang +8

Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two proce…

eess.SP2025

UAV-Enabled Integrated Sensing and Communication in Maritime Emergency Networks

Bohan Li, Jiahao Liu, Junsheng Mu +2

With line-of-sight mode deployment and fast response, unmanned aerial vehicle (UAV), equipped with the cutting-edge integrated sensing and communication (ISAC) technique, is poised…

cs.CV2022

When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition

Bohan Li, Ye Yuan, Dingkang Liang +5

Recently, most handwritten mathematical expression recognition (HMER) methods adopt the encoder-decoder networks, which directly predict the markup sequences from formula images wi…

cs.RO2026

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

GigaWorld Team, Angyuan Ma, Boyuan Wang +24

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow,…

math.QA2023

Spectral flow, twisted modules and MLDE of quasi-lisse vertex algebras

Bohan Li, Hao Li, Wenbin Yan

We calculate the fusion rules among -twisted modules at admissible levels. We derive a series MLDEs for normalized characters of ordinar…

physics.optics2019

An optical Eratosthenes' sieve for large prime numbers

Bohan Li, G. Maltese, J. I. Costa-Filho +2

We report the first experimental demonstration of prime number sieve via linear optics. The prime numbers distribution is encoded in the intensity zeros of the far field produced b…

eess.AS2025

AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions

Yiwei Guo, Bohan Li, Hankun Wang +4

Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensiti…

cs.CL2019

An Adversarial Approach to High-Quality, Sentiment-Controlled Neural Dialogue Generation

Xiang Kong, Bohan Li, Graham Neubig +2

In this work, we propose a method for neural dialogue response generation that allows not only generating semantically reasonable responses according to the dialogue history, but a…

cs.SC2024

A Local Search Algorithm for MaxSMT(LIA)

Xiang He, Bohan Li, Mengyu Zhao +1

MaxSAT modulo theories (MaxSMT) is an important generalization of Satisfiability modulo theories (SMT) with various applications. In this paper, we focus on MaxSMT with the backgro…

cs.HC2026

FIRMED: A Peak-Centered Multimodal Dataset with Fine-Grained Annotation for Emotion Recognition

Hao Tang, Songyun Xie, Xinzhou Xie +4

Traditional video-induced physiological datasets usually rely on whole-trial labels, which introduce temporal label noise in dynamic emotion recognition. We present FIRMED, a peak-…

cs.CV2024

NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation

Baao Xie, Bohan Li, Zequn Zhang +4

3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. T…

cs.SD2026

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

Qibing Bai, Shuai Wang, Yuhan Du +3

Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require natura…

physics.optics2025

Efficient and wavelength-tunable second-harmonic generation towards the green gap

Zhiquan Yuan, Jinhao Ge, Peng Liu +7

Achieving compact and efficient visible laser sources is crucial for a wide range of applications. However traditional semiconductor laser technology faces difficulties in producin…

cs.CV2026

Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

Hu Zhu, Bohan Li, Xianda Guo +5

Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…

cs.CV2025

MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction

Yingshuang Zou, Yikang Ding, Chuanrui Zhang +6

Recent breakthroughs in radiance fields have significantly advanced 3D scene reconstruction and novel view synthesis (NVS) in autonomous driving. Nevertheless, critical limitations…

cs.SE2025

Automated Snippet-Alignment Data Augmentation for Code Translation

Zhiming Zhang, Qingfu Zhu, Xianzhen Luo +3

Code translation aims to translate the code from its source language to the target language and is used in various software development scenarios. Recent developments in Large Lang…

cs.IR2025

MultiRAG: A Knowledge-guided Framework for Mitigating Hallucination in Multi-source Retrieval Augmented Generation

Wenlong Wu, Haofen Wang, Bohan Li +3

Retrieval Augmented Generation (RAG) has emerged as a promising solution to address hallucination issues in Large Language Models (LLMs). However, the integration of multiple retri…

hep-th2023

Superconformal indices of Chern-Simons matter theories

Bohan Li, Dan Xie, WenBin Yan

Gaiotto and Witten found that one can construct 3d Chern-Simons matter theories by using SCFT whose momentum map of global symmetries satisfy specia…

cs.CV2025

MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction

Ziyang Yan, Ruikai Li, Zhiyong Cui +7

Online HD map construction is a fundamental task in autonomous driving systems, aiming to acquire semantic information of map elements around the ego vehicle based on real-time sen…

cs.CV2024

One at a Time: Progressive Multi-step Volumetric Probability Learning for Reliable 3D Scene Perception

Bohan Li, Yasheng Sun, Jingxin Dong +4

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (S…

cs.IR2023

Knowledge Enhancement for Contrastive Multi-Behavior Recommendation

Hongrui Xuan, Yi Liu, Bohan Li +1

A well-designed recommender system can accurately capture the attributes of users and items, reflecting the unique preferences of individuals. Traditional recommendation techniques…

eess.AS2025

A Survey on Speech Large Language Models for Understanding

Jing Peng, Yucheng Wang, Bohan Li +9

Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that…

cs.CV2026

LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding

Hao Li, Minghan Qin, Zhengyu Zou +6

Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from n…

cs.CV2024

Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

Bohan Li, Yasheng Sun, Zhujin Liang +6

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to pre…

cs.CV2025

Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping

Hao Shan, Ruikai Li, Han Jiang +8

As one of the fundamental modules in autonomous driving, online high-definition (HD) maps have attracted significant attention due to their cost-effectiveness and real-time capabil…

cs.IT2023

MDD-Enabled Two-Tier Terahertz Fronthaul in Indoor Industrial Cell-Free Massive MIMO

Bohan Li, Diego Dupleich, Guoqing Xia +4

To make indoor industrial cell-free massive multiple-input multiple-output (CF-mMIMO) networks free from wired fronthaul, this paper studies a multicarrier-division duplex (MDD)-en…

cs.CV2026

MedVAR: Towards Scalable and Efficient Medical Image Generation via Next-scale Autoregressive Prediction

Zhicheng He, Yunpeng Zhao, Junde Wu +5

Medical image generation is pivotal in applications like data augmentation for low-resource clinical tasks and privacy-preserving data sharing. However, developing a scalable gener…

cs.CV2026

Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting

Nan Wang, Yuantao Chen, Lixing Xiao +11

Neural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world scenarios,…

cs.SD2026

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

Wenbin Huang, Yuhang Qiu, Bohan Li +5

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream…

cs.LO2023

Local Search For Satisfiability Modulo Integer Arithmetic Theories

Shaowei Cai, Bohan Li, Xindi Zhang

Satisfiability Modulo Theories (SMT) refers to the problem of deciding the satisfiability of a formula with respect to certain background first order theories. In this paper, we fo…

cs.CV2025

ORV: 4D Occupancy-centric Robot Video Generation

Xiuyu Yang, Bohan Li, Shaocong Xu +9

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, ye…

math.OC2026

Optimal Matching Strategies in Two-sided Markets: A Mean Field Approach

Erhan Bayraktar, Dantong Chu, Bohan Li +1

This paper develops a mean field game framework for dynamic two-sided matching markets, extending existing matching theory by integrating micro-macro dynamics in two-sided environm…

math.RT2024

On some simple orbifold affine VOAs at non-admissible level arising from rank one 4D SCFTs

Tomoyuki Arakawa, Xuanzhong Dai, Justine Fasquel +2

We study the representations of some simple affine vertex algebras at non-admissible level arising from rank one 4D SCFTs. In particular, we classify the irreducible highest weight…

cs.SD2021

AdaSpeech 2: Adaptive Text to Speech with Untranscribed Data

Yuzi Yan, Xu Tan, Bohan Li +4

Text to speech (TTS) is widely used to synthesize personal voice for a target speaker, where a well-trained source TTS model is fine-tuned with few paired adaptation data (speech a…

cs.RO2026

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Zhenduo Shang, Xiyao Liu, Bohan Li +4

Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing met…

cs.SD2026

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Hankun Wang, Bohan Li, Shi Lian +5

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…

cs.LG2026

CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

Yasheng Sun, Bohan Li, Youqi Tao +1

A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind…

cs.SD2026

dots.tts Technical Report

Shi Lian, Changtao Li, Bohan Li +6

We present dots.tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing continuo…

cond-mat.str-el2024

Room-temperature non-volatile optical manipulation of polar order in a charge density wave

Qiaomei Liu, Dong Wu, Tianyi Wu +19

Utilizing ultrafast light-matter interaction to manipulate electronic states of quantum materials is an emerging area of research in condensed matter physics. It has significant im…

cs.CV2023

EMoG: Synthesizing Emotive Co-speech 3D Gesture with Diffusion Model

Lianying Yin, Yijun Wang, Tianyu He +5

Although previous co-speech gesture generation methods are able to synthesize motions in line with speech content, it is still not enough to handle diverse and complicated motion d…