Publications (123)
SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval
Yuxiao Luo, Da Li, Mingjie Zhang +3
LLM-based retrievers have become a fundamental component of modern information retrieval systems. The paradigm of "rewrite-then-retriev" introduces explicit reasoning before retrie…
Kwai Keye-VL Technical Report
Kwai Keye Team, Biao Yang, Bin Wen +57
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form vi…
Prediction Calibration for Generalized Few-shot Semantic Segmentation
Zhihe Lu, Sen He, Da Li +2
Generalized Few-shot Semantic Segmentation (GFSS) aims to segment each image pixel into either base classes with abundant training examples or novel classes with only a handful of…
Interpretability Transfer from Language to Vision via Sparse Autoencoders
Alexey Kravets, Da Li, Chuan Li +2
Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambigui…
A Survey of Link Prediction in N-ary Knowledge Graphs
Jiyao Wei, Saiping Guan, Da Li +3
N-ary Knowledge Graphs (NKGs) are a specialized type of knowledge graph designed to efficiently represent complex real-world facts. Unlike traditional knowledge graphs, where a fac…
Ground-to-UAV sub-Terahertz channel measurement and modeling
Da Li, Peian Li, Jiabiao Zhao +8
Unmanned Aerial Vehicle (UAV) assisted terahertz (THz) wireless communications have been expected to play a vital role in the next generation of wireless networks. UAVs can serve a…
An Electromagnetic-Information-Theory Based Model for Efficient Characterization of MIMO Systems in Complex Space
Ruifeng Li, Da Li, Jinyan Ma +6
It is the pursuit of a multiple-input-multiple-output (MIMO) system to approach and even break the limit of channel capacity. However, it is always a big challenge to efficiently c…
Episodic Training for Domain Generalization
Da Li, Jianshu Zhang, Yongxin Yang +3
Domain generalization (DG) is the challenging and topical problem of learning models that generalize to novel testing domains with different statistics than a set of known training…
Feed-Forward Source-Free Domain Adaptation via Class Prototypes
Ondrej Bohdal, Da Li, Timothy Hospedales
Source-free domain adaptation has become popular because of its practical usefulness and no need to access source data. However, the adaptation process still takes a considerable a…
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
Young D. Kwon, Rui Li, Sijia Li +3
State-of-the-art text-to-image diffusion models (DMs) achieve remarkable quality, yet their massive parameter scale (8-11B) poses significant challenges for inferences on resource-…
Stochastic Gradient Descent in the Viewpoint of Graduated Optimization
Da Li, Jingjing Wu, Qingrun Zhang
Stochastic gradient descent (SGD) method is popular for solving non-convex optimization problems in machine learning. This work investigates SGD from a viewpoint of graduated optim…
Mössbauer spectroscopy study of the magnetostructural and spin-state transitions in the breathing pyrochlore LiFeCrO
Bo Zhang, Wei Ren, Shengyu Yang +9
We report on investigations of the complex magnetostructural and spin-state transitions in the breathing pyrochlore LiFeCrO by means of magnetization, Mössbauer spectr…
Meta Omnium: A Benchmark for General-Purpose Learning-to-Learn
Ondrej Bohdal, Yinbing Tian, Yongshuo Zong +5
Meta-learning and other approaches to few-shot learning are widely studied for image recognition, and are increasingly applied to other vision tasks such as pose estimation and den…
Transmission characteristics of millimeter and sub-terahertz channels through spatially ripple plasma sheath layers
Wenbo Liu, Peian Li, Da Li +2
The propagation of millimeter wave (MMW) and sub-terahertz (THz) signals through plasma sheaths is a critical concern for maintaining communication with hypersonic vehicles, yet th…
Room temperature 2D ferromagnetism in few-layered 1-CrTe
Xingdan Sun, Wanying Li, Xiao Wang +21
Spin-related electronics using two dimensional (2D) van der Waals (vdW) materials as a platform are believed to hold great promise for revolutionizing the next generation spintroni…
SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
Da Li, Jiping Jin, Xuanlong Yu +6
Parametric 3D human models such as SMPL have driven significant advances in human pose and shape estimation, yet their simplified kinematics limit biomechanical realism. The recent…
On-Chip Vectorial Structured Light Manipulation via Inverse Design
Xiaobin Lin, Maoliang Wei, Kunhao Lei +12
On-chip structured light, with potentially infinite complexity, has emerged as a linchpin in the realm of integrated photonics. However, the realization of arbitrarily tailoring a…
Weak-PDE-Net: Discovering Open-Form PDEs via Differentiable Symbolic Networks and Weak Formulation
Xinxin Li, Xingyu Cui, Jin Qi +3
Discovering governing Partial Differential Equations (PDEs) from sparse and noisy data is a challenging issue in data-driven scientific computing. Conventional sparse regression me…
The unexpected binding and superconductivity in SbH4 at high pressure
Yanbin Ma, Defang Duan, Da Li +7
The semimetal antimony (Sb) element doped into hydrogen has been performed theoretically to explored high-pressure crystal structure and superconductivity of antimony hydrides. The…
Dual-Function Beamforming Design For Multi-Target Localization and Reliable Communications
Bo Tang, Da Li, Wenjun Wu +3
This paper investigates the transmit beamforming design for multiple-input multiple-output systems to support both multi-target localization and multi-user communications. To enhan…
LifeIR at the NTCIR-18 Lifelog-6 Task
Jiahan Chen, Da Li, Keping Bi
In recent years, sharing lifelogs recorded through wearable devices such as sports watches and GoPros, has gained significant popularity. Lifelogs involve various types of informat…
ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
Ruchika Chavhan, Da Li, Timothy Hospedales
While large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generat…
MMDG-Bench: A Benchmark for Multimodal Domain Generalization
Qianshan Zhan, Qian Wang, Da Li +2
Multi-modal Domain Generalization (MMDG) seeks to leverage complementary modalities to enhance model robustness on unseen domains. Despite extensive progress in Multi-modal Learnin…
AMS_ADRN at SemEval-2022 Task 5: A Suitable Image-text Multimodal Joint Modeling Method for Multi-task Misogyny Identification
Da Li, Ming Yi, Yukai He
Women are influential online, especially in image-based social media such as Twitter and Instagram. However, many in the network environment contain gender discrimination and aggre…
Neural Fine-Tuning Search for Few-Shot Learning
Panagiotis Eustratiadis, Åukasz Dudziak, Da Li +1
In few-shot recognition, a classifier that has been trained on one set of classes is required to rapidly adapt and generalize to a disjoint, novel set of classes. To that end, rece…
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
Wenqing Wang, Da Li, Xiatian Zhu +1
Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adapt…
Generative Model Based Noise Robust Training for Unsupervised Domain Adaptation
Zhongying Deng, Da Li, Junjun He +2
Target domain pseudo-labelling has shown effectiveness in unsupervised domain adaptation (UDA). However, pseudo-labels of unlabeled target domain data are inevitably noisy due to t…
Breaking the Degrees-of-Freedom Limit of Holographic MIMO Communications: A 3-D Antenna Array Topology
Shuai S. A. Yuan, Jie Wu, Hongjing Xu +9
The performance of holographic multiple-input multiple-output (MIMO) communications, employing two-dimensional (2-D) planar antenna arrays, is typically compromised by finite degre…
High-speed surface-property recognition by 140-GHz frequency
Jiacheng Liu, Da Li, Guohao Liu +5
In the field of integrated sensing and communication, there's a growing need for advanced environmental perception. The terahertz (THz) frequency band, significant for ultra-high-s…
Measurement and Modeling on Terahertz Channels in Rain
Peian Li, Wenbo Liu, Jiacheng Liu +7
The Terahertz (THz) frequency band offers a wide range of bandwidths, from tens to hundreds of gigahertz (GHz) and also supports data speeds of several terabits per second (Tbps).…
Attacking Adversarial Defences by Smoothing the Loss Landscape
Panagiotis Eustratiadis, Henry Gouk, Da Li +1
This paper investigates a family of methods for defending against adversarial attacks that owe part of their success to creating a noisy, discontinuous, or otherwise rugged loss la…
Kwai Keye-VL 1.5 Technical Report
Biao Yang, Bin Wen, Boyang Ding +58
In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…
Multi-task Pre-training Language Model for Semantic Network Completion
Da Li, Sen Yang, Kele Xu +3
Semantic networks, such as the knowledge graph, can represent the knowledge leveraging the graph structure. Although the knowledge graph shows promising values in natural language…
EditSR: Enhancing Neural Symbolic Regression via Edit-based Rectification
Da Li, Xinxin Li, Xingyu Cui +3
Neural symbolic regression models improve inference efficiency by shifting structural search to pretraining, but their one-pass autoregressive decoding is prone to error accumulati…
Urban Neural Surface Reconstruction from Constrained Sparse Aerial Imagery with 3D SAR Fusion
Da Li, Chen Yao, Tong Mao +2
Neural surface reconstruction (NSR) has recently shown strong potential for urban 3D reconstruction from multi-view aerial imagery. However, existing NSR methods often suffer from…
Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable Style
Fengyin Lin, Mingkang Li, Da Li +3
This paper studies the problem of zero-short sketch-based image retrieval (ZS-SBIR), however with two significant differentiators to prior art (i) we tackle all variants (inter-cat…
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
Da Li, Keping Bi, Jiafeng Guo +1
Table retrieval, essential for accessing information through tabular data, is less explored compared to text retrieval. The row/column structure and distinct fields of tables (incl…
Modulated non-collinear magnetic structure of (CoFe)NbO as revealed by Mössbauer spectroscopy
Bo Zhang, Qifeng Kuang, Hua Pang +4
In this work, we present detailed Fe Mössbauer spectroscopy investigations of (CoFe)NbO compound to study its possible magnetic s…
Intermittent Josephson effect with feedback voltage and temperature oscillations in graphite-coated nanocapsules with superconducting TaC core
Dianyu Geng, Zhenhua Wang, Da Li +2
An intermittent Josephson effect in the form of voltage and temperature oscillations in the voltage - current curves near 2 K is observed in pellets consisting of superconducting T…
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
Lihao Liu, Yan Wang, Biao Yang +10
Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct applic…
Learning to Generalize: Meta-Learning for Domain Generalization
Da Li, Yongxin Yang, Yi-Zhe Song +1
Domain shift refers to the well known problem that a model trained in one source domain performs poorly when applied to a target domain with different statistics. {Domain Generaliz…
SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models
Ruolin Yang, Da Li, Honggang Zhang +1
Sketching is a uniquely human tool for expressing ideas and creativity. The animation of sketches infuses life into these static drawings, opening a new dimension for designers. An…
Explore of exfoliable multifunctional high-k two-dimensional oxides
Yue Hu, Jingwen Jiang, Peng Zhang +5
As the continuing down-scaling of field-effect transistors (FETs) in more-than-Moore integrated circuits, finding new functional two-dimensional (2D) materials with a higher dielec…
On the Limitations of General Purpose Domain Generalisation Methods
Henry Gouk, Ondrej Bohdal, Da Li +1
We investigate the fundamental performance limitations of learning algorithms in several Domain Generalisation (DG) settings. Motivated by the difficulty with which previously prop…
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
Da Li, Yuxiao Luo, Keping Bi +7
Multimodal Large Language Models advance multimodal representation learning by acquiring transferable semantic embeddings, thereby substantially enhancing performance across a rang…
SkinningGS: Editable Dynamic Human Scene Reconstruction Using Gaussian Splatting Based on a Skinning Model
Da Li, Donggang Jia, Markus Hadwiger +1
Reconstructing an interactive human avatar and the background from a monocular video of a dynamic human scene is highly challenging. In this work we adopt a strategy of point cloud…
Dynamic Instance Domain Adaptation
Zhongying Deng, Kaiyang Zhou, Da Li +3
Most existing studies on unsupervised domain adaptation (UDA) assume that each domain's training samples come with domain labels (e.g., painting, photo). Samples from each domain a…
Application of an unbalanced optimal transport distance and a mixed L1/Wasserstein distance to full waveform inversion
Da Li, Michael P. Lamoureux, Wenyuan Liao
Full waveform inversion (FWI) is an important and popular technique in subsurface earth property estimation. However, using the least-squares norm in the misfit function often lead…
Online Meta-Learning for Multi-Source and Semi-Supervised Domain Adaptation
Da Li, Timothy Hospedales
Domain adaptation (DA) is the topical problem of adapting models from labelled source datasets so that they perform well on target datasets where only unlabelled or partially label…
Phase diagram and superconductivity of polonium hydrides under high pressure
Yunxian Liu, Defang Duan, Fubo Tian +8
High pressure structures, phase diagram and superconductivity of polonium hydrides have been systematically investigated through the first-principles calculations based on the dens…
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
Zhen Wang, Da Li, Yulin Su +3
Logo embedding models convert the product logos in images into vectors, enabling their utilization for logo recognition and detection within e-commerce platforms. This facilitates…
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework
Zhuo Zhi, Chen Feng, Adam Daneshmend +6
Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic fr…
Terahertz channel modeling based on surface sensing characteristics
Jiayuan Cui, Da Li, Jiabiao Zhao +7
The dielectric properties of environmental surfaces, including walls, floors and the ground, etc., play a crucial role in shaping the accuracy of terahertz (THz) channel modeling,…
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
Zeyuan Wang, Da Li, Yulin Chen +4
We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with…
Electric-field control of magnetism in few-layered van der Waals magnet
Zhi Wang, Tong-Yao Zhang, Mei Ding +16
Manipulating quantum state via electrostatic gating has been intriguing for many model systems in nanoelectronics. When it comes to the question of controlling the electron spins,…
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Zixu Cheng, Da Li, Jian Hu +4
Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language…
Relative Entropy-Based Waveform Optimization for Rician Target Detection with Dual-Function Radar Communication Systems
Xuyang Wang, Bo Tang, Wenjun Wu +1
In this paper, we consider waveform design for dualfunction radar-communication systems based on multiple-inputmultiple-out arrays. To achieve better Rician target detection perfor…
Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a Difference
Shell Xu Hu, Da Li, Jan Stühmer +2
Few-shot learning (FSL) is an important and topical problem in computer vision that has motivated extensive research into numerous methods spanning from sophisticated meta-learning…
Deep Factorised Inverse-Sketching
Kaiyue Pang, Da Li, Jifei Song +3
Modelling human free-hand sketches has become topical recently, driven by practical applications such as fine-grained sketch based image retrieval (FG-SBIR). Sketches are clearly r…
Fast Integral Histogram Computations on GPU for Real-Time Video Analytics
Mahdieh Poostchi, Kannappan Palaniappan, Da Li +3
In many Multimedia content analytics frameworks feature likelihood maps represented as histograms play a critical role in the overall algorithm. Integral histograms provide an effi…
Better Practices for Domain Adaptation
Linus Ericsson, Da Li, Timothy M. Hospedales
Distribution shifts are all too common in real-world applications of machine learning. Domain adaptation (DA) aims to address this by providing various frameworks for adapting mode…
PushDualGen: Enabling LLMs to Generate Semantic IDs with Interpretable Copy for Industrial Push Recommendation
Manjia Lin, Da Li, Yan Wang +9
Push recommendation in KuaiShou proactively delivers personalized content to nearly one billion users to facilitate their engagement. Recently, generative recommendation has achiev…
Love in Action: Gamifying Public Video Cameras for Fostering Social Relationships in Real World
Zhang Zhang, Da Li, Geng Wu +3
In this paper, we create "Love in Action" (LIA), a body language-based social game utilizing video cameras installed in public spaces to enhance social relationships in real-world.…
FedL2P: Federated Learning to Personalize
Royson Lee, Minyoung Kim, Da Li +4
Federated learning (FL) research has made progress in developing algorithms for distributed learning of global models, as well as algorithms for local personalization of those comm…
Bridging Queries and Tables through Entities in Table Retrieval
Da Li, Keping Bi, Jiafeng Guo +1
Table retrieval is essential for accessing information stored in structured tabular formats; however, it remains less explored than text retrieval. The content of the table primari…
Learning to Augment via Implicit Differentiation for Domain Generalization
Tingwei Wang, Da Li, Kaiyang Zhou +2
Machine learning models are intrinsically vulnerable to domain shift between training and testing data, resulting in poor performance in novel domains. Domain generalization (DG) a…
Impact of snowfall on terahertz channel performance: measurement and modeling insights
Guohao Liu, Xiangkun He, Jiabiao Zhao +5
In the evolving domain of wireless communication, the investigation on terahertz (THz) frequency spectrum, spanning 0.1 to 10 THz, has become a critical focus for advancing ultra-h…
Unveiling the spontaneous conversion of layered MAX phases to 2D MXenes
Tao Hu, Shihao Zhu, Zhaojin Li +4
Topochemically transforming layered non-van der Waals solid into two dimensional (2D) materials involves selective etching reactions with atomic precision. The element-specific, st…
Terahertz channel power and BER performance in rain
Yuheng Song, Jiayuan Cui, Guohao Liu +9
Terahertz (THz) communications have emerged as a promising technology for 6G networks due to their potential for achieving terabit-per-second data rates. However, the impact of rai…
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5
The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…
Weight-Covariance Alignment for Adversarially Robust Neural Networks
Panagiotis Eustratiadis, Henry Gouk, Da Li +1
Stochastic Neural Networks (SNNs) that inject noise into their hidden layers have recently been shown to achieve strong robustness against adversarial attacks. However, existing SN…
Efficient and Stable Finite Difference Modelling of Acoustic Wave Propagation in Variable-density Media
Da Li, Keran Li, Wenyuan Liao
In this paper, we consider the development and analysis of a new explicit compact high-order finite difference scheme for acoustic wave equation formulated in divergence form, whic…
Pressure-induced decomposition of solid hydrogen sulfide
Defang Duan, Xiaoli Huang, Fubo Tian +7
Solid hydrogen sulfide is well known as a typical molecular crystal but its stability under pressure is still under debate. Particularly, Eremets et al. found the high pressure sup…
Compiler-Assisted Workload Consolidation For Efficient Dynamic Parallelism on GPU
Hancheng Wu, Da Li, Michela Becchi
GPUs have been widely used to accelerate computations exhibiting simple patterns of parallelism - such as flat or two-level parallelism - and a degree of parallelism that can be st…
Eavesdropping risk evaluation for non-line-of-sight terahertz channels by metallic wavy surface in rain
Peian Li, Wenbo Liu, Da Li +4
Non-line-of-sight (NLOS) data transmission through surface reflection is pivotal for enhancing the reach and efficiency of terahertz (THz) communication systems. However, this inno…
Broken mirror symmetry tuned topological transport in PbTe/SnTe heterostructures
Feng Wei, Chieh-Wen Liu, Da Li +6
The tunability of topological surface states and controllable opening of the Dirac gap are of great importance to the application of topological materials. In topological crystalli…
Observation of non-Hermitian topological disclination states and charge fractionalization
Ruifeng Li, Rimi Banerjee, Subhaskar Mandal +8
There has been significant interest in exploring topological disclination states, which effectively probe the band topology of the host material beyond the conventional bulk-edge c…
Label Calibration for Semantic Segmentation Under Domain Shift
Ondrej Bohdal, Da Li, Timothy Hospedales
Performance of a pre-trained semantic segmentation model is likely to substantially decrease on data from a new domain. We show a pre-trained model can be adapted to unlabelled tar…
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning
Jiahan Chen, Da Li, Hengran Zhang +6
Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classi…
Decomposition of solid hydrogen bromide at high pressure
Defang Duan, Fubo Tian, Xiaoli Huang +6
The stability of different stoichiometric HBr (=1-7) compounds under pressure are extensively studied using density functional theory calculations. Five new energetically st…
UniSymNet: A Unified Symbolic Network Guided by Transformer
Xinxin Li, Juan Zhang, Da Li +3
Symbolic Regression (SR) is a powerful technique for automatically discovering mathematical expressions from input data. Mainstream SR algorithms search for the optimal symbolic tr…
Sequential Learning for Domain Generalization
Da Li, Yongxin Yang, Yi-Zhe Song +1
In this paper we propose a sequential learning framework for Domain Generalization (DG), the problem of training a model that is robust to domain shift by design. Various DG approa…
Adapt PointFormer: 3D Point Cloud Analysis via Adapting 2D Visual Transformers
Mengke Li, Da Li, Guoqing Yang +2
Pre-trained large-scale models have exhibited remarkable efficacy in computer vision, particularly for 2D image analysis. However, when it comes to 3D point clouds, the constrained…
Robust Target Training for Multi-Source Domain Adaptation
Zhongying Deng, Da Li, Yi-Zhe Song +1
Given multiple labeled source domains and a single target domain, most existing multi-source domain adaptation (MSDA) models are trained on data from all domains jointly in one ste…
UAV-Assisted Weather Radar Calibration: A Theoretical Model for Wind Influence on Metal Sphere Reflectivity
Jiabiao Zhao, Da Li, Jiayuan Cui +2
The calibration of weather radar for detecting meteorological phenomena has advanced rapidly, aiming to enhance accuracy. Utilizing an unmanned aerial vehicle (UAV) equipped with a…
Chromium-Induced Ferromagnetism with Perpendicular Anisotropy in Topological Crystalline Insulator SnTe (111) Thin Films
Fei Wang, Hongrui Zhang, Jue Jiang +8
Topological crystalline insulator (TCI) is a recently-discovered topological phase of matter. It possesses multiple Dirac surface states, which are protected by the crystal symmetr…
Domain Generalisation via Domain Adaptation: An Adversarial Fourier Amplitude Approach
Minyoung Kim, Da Li, Timothy Hospedales
We tackle the domain generalisation (DG) problem by posing it as a domain adaptation (DA) task where we adversarially synthesise the worst-case target domain and adapt a model to t…
Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation
Yu Zhong, Zihao Zhang, Rui Zhang +9
Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of har…
Measurement and Modeling on Terahertz Channel Propagation Through Vegetation
Jiayuan Cui, Yuheng Song, He Jiang +11
The terahertz band offers promising opportunities for high-capacity wireless communications but faces significant challenges from vegetation-induced channel impairments. This artic…
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
Zeyuan Wang, Da Li, Yulin Chen +6
Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but str…
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
Zechang Xiong, Da Li, Kexin Tang +3
Multimodal models often converge to a dominant-modality solution, in which a stronger, faster-converging modality overshadows weaker ones. This modality imbalance causes suboptimal…
Co-Design for Spectral Coexistence between RIS-aided MIMO Radar and MIMO Communication Systems
Da Li, Bo Tang, Xuyang Wang +2
Reconfigurable intelligent surface (RIS) refers to a signal reflection surface containing a large number of low-cost passive reflecting elements. RIS can improve the performance of…
STSC-SNN: Spatio-Temporal Synaptic Connection with Temporal Convolution and Attention for Spiking Neural Networks
Chengting Yu, Zheming Gu, Da Li +3
Spiking Neural Networks (SNNs), as one of the algorithmic models in neuromorphic computing, have gained a great deal of research attention owing to temporal information processing…
A Large-scale Distributed Video Parsing and Evaluation Platform
Kai Yu, Yang Zhou, Da Li +2
Visual surveillance systems have become one of the largest data sources of Big Visual Data in real world. However, existing systems for video analysis still lack the ability to han…
RaRa Clipper: A Clipper for Gaussian Splatting Based on Ray Tracer and Rasterizer
Da Li, Donggang Jia, Yousef Rajeh +2
With the advancement of Gaussian Splatting techniques, a growing number of datasets based on this representation have been developed. However, performing accurate and efficient cli…
Sketch-based Video Object Segmentation: Benchmark and Analysis
Ruolin Yang, Da Li, Conghui Hu +3
Reference-based video object segmentation is an emerging topic which aims to segment the corresponding target object in each video frame referred by a given reference, such as a la…
Deep Learning-based Cross-modal Reconstruction of Vehicle Target from Sparse 3D SAR Image
Da Li, Guoqiang Zhao, Chen Yao +4
Three-dimensional synthetic aperture radar (3D SAR) is an advanced active microwave imaging technology widely utilized in remote sensing area. To achieve high-resolution 3D imaging…
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
Haochen Chang, Pengfei Ren, Buyuan Zhang +6
Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestu…
High-pressure structures and superconductivity of bismuth hydrides
Yanbin Ma, Defang Duan, Da Li +7
We have systematically searched for the ground state structures of bismuth hydrides based on evolutionary algorithm method and particle swarm optimization algorithm method. Given o…
Extraction of n = 0 pick-up by locked mode detectors based on neural networks in J-TEXT
Chengshuo Shen, Jianchao Li, Yonghua Ding +10
Measurement of locked mode (LM) is important for the physical research of Magnetohydrodynamic (MHD) instabilities and plasma disruption. The n = 0 pick-up need to be extracted and…