Publications (439)
Compact superconducting microwave resonators based on Al-AlOx-Al capacitor
Julia Zotova, Rui Wang, Alexander Semenov +5
We address the scaling-up problem for superconducting quantum circuits by using lumped-element resonators based on an alternative fabrication method of aluminum -- aluminum oxide -…
Implicit Neural Image Field for Biological Microscopy Image Compression
Gaole Dai, Cheng-Ching Tseng, Qingpo Wuwu +9
The rapid pace of innovation in biological microscopy imaging has led to large images, putting pressure on data storage and impeding efficient sharing, management, and visualizatio…
DualX-VSR: Dual Axial SpatialTemporal Transformer for Real-World Video Super-Resolution without Motion Compensation
Shuo Cao, Yihao Liu, Xiaohui Li +3
Transformer-based models like ViViT and TimeSformer have advanced video understanding by effectively modeling spatiotemporal dependencies. Recent video generation models, such as S…
A Syntax-Guided Multi-Task Learning Approach for Turducken-Style Code Generation
Guang Yang, Yu Zhou, Xiang Chen +4
Due to the development of pre-trained language models, automated code generation techniques have shown great promise in recent years. However, the generated code is difficult to me…
Waveform and Filter Design for Integrated Sensing and Communication Against Signal-dependent Modulated Jamming
Yu Zhou, Qiao Shi, Zhengchun Zhou +2
This paper focuses on an integrated sensing and communication (ISAC) system in the presence of signal-dependent modulated jamming (SDMJ). Our goal is to suppress jamming while carr…
Beyond Instance Discrimination: Relation-aware Contrastive Self-supervised Learning
Yifei Zhang, Chang Liu, Yu Zhou +3
Contrastive self-supervised learning (CSL) based on instance discrimination typically attracts positive samples while repelling negatives to learn representations with pre-defined…
Inverse Kinematics on Guiding Vector Fields for Robot Path Following
Yu Zhou, Jesús Bautista, Weijia Yao +1
Inverse kinematics is a fundamental technique for motion and positioning control in robotics, typically applied to end-effectors. In this paper, we extend the concept of inverse ki…
Hard Label Black Box Node Injection Attack on Graph Neural Networks
Yu Zhou, Zihao Dong, Guofeng Zhang +1
While graph neural networks have achieved state-of-the-art performances in many real-world tasks including graph classification and node classification, recent works have demonstra…
Neural Abstract Style Transfer for Chinese Traditional Painting
Bo Li, Caiming Xiong, Tianfu Wu +3
Chinese traditional painting is one of the most historical artworks in the world. It is very popular in Eastern and Southeast Asia due to being aesthetically appealing. Compared wi…
Interplay between defects and the non-Hermitian skin effect
Yin Huang, Wenna Zhang, Yu Zhou +3
The non-Hermitian skin effect (NHSE) is an intriguing phenomenon in which an extensive number of bulk eigenstates localize at the boundaries of a non-Hermitian system with non-reci…
Observation and manipulation of quantum interference in a superconducting Kerr parametric oscillator
Daisuke Iyama, Takahiko Kamiya, Shiori Fujii +9
Quantum tunneling is the phenomenon that makes superconducting circuits "quantum". Recently, there has been a renewed interest in using quantum tunneling in phase space of a Kerr p…
A Highly Efficient Cross-matching Scheme using Learned Index Structure
Phu-Minh Lam, Dongwei Fan, Hongbo Wei +6
Spatial data fusion is a bottleneck when it meets the scale of 10 billion records. Cross-matching celestial catalogs is just one example of this. To challenge this, we present a fr…
Progressive Cluster Purification for Unsupervised Feature Learning
Yifei Zhang, Chang Liu, Yu Zhou +3
In unsupervised feature learning, sample specificity based methods ignore the inter-class information, which deteriorates the discriminative capability of representation models. Cl…
InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
Chenchen Shou, Guyue Liu, Hao Nie +11
Scaling Large Language Model (LLM) training relies on multi-dimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like T…
Self-Training for Domain Adaptive Scene Text Detection
Yudi Chen, Wei Wang, Yu Zhou +3
Though deep learning based scene text detection has achieved great progress, well-trained detectors suffer from severe performance degradation for different domains. In general, a…
CVD nanodiamonds with non-blinking, near transform-limited linewidths emitters
Ke Li, Yu Zhou, Abdullah Rasmita +2
Near transform-limited single photon sources are required for perfect photon indistinguishability in quantum networks. Having such sources in nanodiamonds is particularly important…
Memory Remedy: An AI-Enhanced Interactive Story Exploring Human-Robot Interaction and Companionship
Lei Han, Yu Zhou, Qiongyan Chen +1
We present our approach to using AI-generated content (AIGC) and multiple media to develop an immersive, game-based, interactive story experience. The narrative of the story, "Memo…
Orochi: Versatile Biomedical Image Processor
Gaole Dai, Chenghao Zhou, Yu Zhou +6
Deep learning has emerged as a pivotal tool for accelerating research in the life sciences, with the low-level processing of biomedical images (e.g., registration, fusion, restorat…
Deterministic and Scalable Coupling of Single 4H-SiC Spin Defects into Bullseye Cavities
Tongyuan Bao, Qi Luo, Ailun Yi +5
Silicon carbide (SiC) has attracted significant attention as a promising quantum material due to its ability to host long-lived, optically addressable color centers with solid-stat…
GEMs: Breaking the Long-Sequence Barrier in Generative Recommendation with a Multi-Stream Decoder
Yu Zhou, Chengcheng Guo, Kuo Cai +6
While generative recommendations (GR) possess strong sequential reasoning capabilities, they face significant challenges when processing extremely long user behavior sequences: the…
A Novel Strategy to Strengthen Directionally Solidified Superalloy Through Grain Boundary Simplified Design
Yunpeng Fan, Xinbao Zhao, Yu Zhou +4
Conventional strategies for enhancing creep resistance often rely on grain boundary strengthening, yet this approach can inadvertently promote premature grain boundary fracture. Th…
Recurrent Meta-Structure for Robust Similarity Measure in Heterogeneous Information Networks
Yu Zhou, Jianbin Huang, Heli Sun +1
Similarity measure as a fundamental task in heterogeneous information network analysis has been applied to many areas, e.g., product recommendation, clustering and Web search. Most…
Endomorphism algebras of 2-term silting complexes
Aslak Bakke Buan, Yu Zhou
We study possible values of the global dimension of endomorphism algebras of 2-term silting complexes. We show that for any algebra whose global dimension $\mathop{\rm gl. dim}…
STEP3-VL-10B Technical Report
Ailin Huang, Chengyuan Yao, Chunrui Han +90
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…
Position-sensitive spectral splitting with a plasmonic nanowire on silicon chip
Qing Hu, Di-Hu Xu, Yu Zhou +6
On-chip nanophotonics serves as the foundation for the new generation of information technology, but it is challenged by the diffraction limit of light. With the capabilities of co…
A Diffusion Model Translator for Efficient Image-to-Image Translation
Mengfei Xia, Yu Zhou, Ran Yi +2
Applying diffusion models to image-to-image translation (I2I) has recently received increasing attention due to its practical applications. Previous attempts inject information fro…
Expert Training: Task Hardness Aware Meta-Learning for Few-Shot Classification
Yucan Zhou, Yu Wang, Jianfei Cai +3
Deep neural networks are highly effective when a large number of labeled samples are available but fail with few-shot classification tasks. Recently, meta-learning methods have rec…
Accurate six-band nearest-neighbor tight-binding model for the pi-bands of bulk graphene and graphene nanoribbons
Timothy B. Boykin, Mathieu Luisier, Gerhard Klimeck +4
Accurate modeling of the pi-bands of armchair graphene nanoribbons (AGNRs) requires correctly reproducing asymmetries in the bulk graphene bands as well as providing a realistic mo…
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
Daiqing Wu, Xuan Zhang, Dongbao Yang +7
The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate tex…
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…
Effective Action and Gravitational Pair Production in (A)dS Spacetime
Yu Zhou, Hai-Qing Zhang
We compute the effective action for a massive scalar field in (A)dS spacetime using the Euclidean heat kernel method. We highlight that in even-dimensional dS spacetimes, the effec…
An Automatic Method for Generating Symbolic Expressions of Zernike Circular Polynomials
Hong-Yan Zhang, Yu Zhou, Fu-Yun Li
Zernike circular polynomials (ZCP) play a significant role in optics engineering. The symbolic expressions for ZCP are valuable for theoretic analysis and engineering designs. Howe…
Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection
Xugong Qin, Yu Zhou, Youhui Guo +5
Due to the large success in object detection and instance segmentation, Mask R-CNN attracts great attention and is widely adopted as a strong baseline for arbitrary-shaped scene te…
AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenarios
Ziming Huang, Xurui Li, Haotian Liu +3
Recently, multi-class anomaly classification has garnered increasing attention. Previous methods directly cluster anomalies but often struggle due to the lack of anomaly-prior know…
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +7
End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and render…
Probing ALP-Photon Mixing with High-Resolution X-ray Spectroscopy
Yu Zhou, Jiejia Liu, Volodymyr Takhistov +1
Axion-like particles (ALPs) provide a compelling avenue for exploring physics beyond the Standard Model. In astrophysical magnetized plasmas an ALP-photon coupling induce…
Exploring Instance Relations for Unsupervised Feature Embedding
Yifei Zhang, Yu Zhou, Weiping Wang
Despite the great progress achieved in unsupervised feature embedding, existing contrastive learning methods typically pursue view-invariant representations through attracting posi…
Entanglement Model for Mode-Pairing Quantum Key Distribution
Yi-Fei Lu, Yang Wang, Hong-Wei Li +9
Mode-pairing (MP) quantum key distribution (QKD) eliminates the requirements of phase locking and phase tracking compared with twin-field (TF) QKD while still surpassing the fundam…
Deutsch's algorithm with topological charges of optical vortices via non-degenerate four-wave mixing
Mingtao Cao, Liang Han, Ruifeng Liu +8
We propose a scheme to implement the Deutsch's algorithm through non-degenerate four-wave mixing process. By employing photon topological charges of optical vortices, we demonstrat…
All-optical naked-eye ghost imaging
Gao Wang, Huaibin Zheng, Yu Zhou +6
Ghost imaging is usually based on optoelectronic process and eletronic computing. We here propose a new ghost imaging scheme, which avoids any optoelectronic or electronic process.…
Cluster categories for marked surfaces: punctured case
Yu Qiu, Yu Zhou
We study the cluster categories arising from marked surfaces (with punctures and non-empty boundaries). By constructing skewed-gentle algebras, we show that there is a bijection be…
Highly-efficient spintronic terahertz emitter enabled by metal-dielectric photonic crystal
Zheng Feng, Rui Yu, Yu Zhou +9
Spintronic terahertz (THz) emitter provides the advantages such as apparently broader spectrum, significantly lower cost, and more flexibility in compared with the commercial THz e…
Two-photon superbunching effect of broadband chaotic stationary light at femtosecond timescale based on cascaded Michelson interferometer
Sheng Luo, Yu Zhou, Huaibin Zheng +7
It is challenging for observing superbunching effect with true chaotic light, here we propose and demonstrate a method to achieve superbunching effect of the degree of second-order…
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
Daiqing Wu, Dongbao Yang, Sicheng Zhao +2
The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of m…
VidText: Towards Comprehensive Evaluation for Video Text Understanding
Zhoufaran Yang, Yan Shu, Jing Wang +8
Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. Howeve…
A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation
Ruoxin Wang, Tianyi Tang, Haiming Du +7
Brain tumor segmentation models have aided diagnosis in recent years. However, they face MRI complexity and variability challenges, including irregular shapes and unclear boundarie…
Leverage Knowledge Graph and Large Language Model for Law Article Recommendation: A Case Study of Chinese Criminal Law
Yongming Chen, Miner Chen, Ye Zhu +7
Judicial efficiency is critical to social stability. However, in many countries worldwide, grassroots courts face substantial case backlogs, and judicial decisions remain heavily d…
Behavior-aware Service Access Control Mechanism using Security Policy Monitoring for SOA Systems
Yunfei Meng, Zhiqiu Huang, Senzhang Wang +3
Service-oriented architecture (SOA) system has been widely utilized at many present business areas. However, SOA system is loosely coupled with multiple services and lacks the rele…
PromptDLA: A Domain-aware Prompt Document Layout Analysis Framework with Descriptive Knowledge as a Cue
Zirui Zhang, Yaping Zhang, Lu Xiang +4
Document Layout Analysis (DLA) is crucial for document artificial intelligence and has recently received increasing attention, resulting in an influx of large-scale public DLA data…
Classification and Generation of Light Sources Using Gamma Fitting
Shuanghao Zhang, Huaibin Zheng, Gao Wang +6
In general, the typical approach to discriminate antibunching, bunching or superbunching categories make use of calculating the second-order coherence function of l…
Class-Agnostic Region-of-Interest Matching in Document Images
Demin Zhang, Jiahao Lyu, Zhijie Shen +1
Document understanding and analysis have received a lot of attention due to their widespread application. However, existing document analysis solutions, such as document layout ana…
Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow
Jiangling Zhang, Shuxuan Gao, Zeyu Chen +2
Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit spec…
A Canonical Structure for Constructing Projected First-Order Algorithms With Delayed Feedback
Mengmou Li, Yu Zhou, Xun Shen +1
This work introduces a canonical structure for a broad class of unconstrained first-order algorithms that admit a Lur'e representation, including systems with relative degree great…
High Order Expansion Method for Kuiper's Statistic in Goodness-of-fit Test
Hong-Yan Zhang, Zhi-Qiang Feng, Haoting Liu +2
Kuiper's statistic, a measure for comparing the difference of ideal distribution and empirical distribution, is of great significance in the goodness-of-fit test. However, Ku…
M3D: Dual-Stream Selective State Spaces and Depth-Driven Framework for High-Fidelity Single-View 3D Reconstruction
Luoxi Zhang, Pragyan Shrestha, Yu Zhou +2
The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neur…
TransCrowd: weakly-supervised crowd counting with transformers
Dingkang Liang, Xiwu Chen, Wei Xu +2
The mainstream crowd counting methods usually utilize the convolution neural network (CNN) to regress a density map, requiring point-level annotations. However, annotating each per…
High visibility temporal ghost imaging with classical light
Jianbin Liu, Jingjing Wang, Hui Chen +5
High visibility temporal ghost imaging with classical light is possible when superbunching pseudothermal light is employed. In the numerical simulation, the visibility of temporal…
DCA: Dividing and Conquering Amnesia in Incremental Object Detection
Aoting Zhang, Dongbao Yang, Chang Liu +3
Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous clas…
Finite-Time Optimization via Scaled Gradient-Momentum Flows
Yu Zhou, Mengmou Li, Masaaki Nagahara
In this paper, we develop a scaled gradient-momentum framework for continuous-time optimization that achieves global finite-time convergence. A state-dependent scaling mechanism is…
Geocoronal Solar Wind Charge Exchange Process Associated with the 2006-December-13 Coronal Mass Ejection Event
Yu Zhou, Noriko Y. Yamasaki, Shin Toriumi +1
We report the discovery of a geocoronal solar wind charge exchange (SWCX) event corresponding to the well-known 2006 December 13th coronal mass ejection (CME) event. Strong evidenc…
Which and Where to Focus: A Simple yet Accurate Framework for Arbitrary-Shaped Nearby Text Detection in Scene Images
Youhui Guo, Yu Zhou, Xugong Qin +1
Scene text detection has drawn the close attention of researchers. Though many methods have been proposed for horizontal and oriented texts, previous methods may not perform well w…
Fermionic ghost imaging
Jianbin Liu, Yu Zhou, Huaibin Zheng +3
Ghost imaging with thermal fermions is calculated based on two-particle interference in Feynman's path integral theory. It is found that ghost imaging with thermal fermions can be…
Semi-invisible Hyperon Decays in the Effective Lagrangian Approach
Lai Jiang, Ye Xing, Yu Zhou +1
We systematically investigate the semi-invisible decays of hyperons (hyperon invisible()) in the Mesogenesis mechanism by the effective Lagrangian approach. Th…
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Shuo Cao, Nan Ma, Jiayang Li +12
The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehens…
Random Delayed-Choice Quantum Eraser via Two-Photon Imaging
Giuliano Scarcelli, Yu Zhou, Yanhua Shih
We report on a delayed-choice quantum eraser experiment based on a two-photon imaging scheme using entangled photon pairs. After the detection of a photon which passed through a do…
Advancing Transformer's Capabilities in Commonsense Reasoning
Yu Zhou, Yunqiu Han, Hanyu Zhou +1
Recent advances in general purpose pre-trained language models have shown great potential in commonsense reasoning. However, current works still perform poorly on standard commonse…
Artificial intelligence control of a turbulent jet
Yu Zhou, Dewei Fan, Bingfu Zhang +2
An artificial intelligence (AI) control system is developed to maximize the mixing rate of a turbulent jet. This system comprises six independently operated unsteady minijet actuat…
A Method of Measuring TES Complex ETF Response in Frequency-domain Multiplexed Readout by Single Sideband Power Modulation
Yu Zhou, Tijmen de Haan, Hiroki Akamatsu +5
The digital frequency domain multiplexing (DfMux) technique is widely used for astrophysical instruments with large detector arrays. Detailed detector characterization is required…
CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization
Min Xiao, Junnan Zhu, Haitao Lin +2
Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear. Existing multimodal summarization approaches focus on designing t…
Semibricks and wide subcategories in extended module categories
Esha Gupta, Yu Zhou
For , we define semibricks and wide subcategories in the -extended hearts of bounded -structures on a triangulated category. We show that these semibricks are in bij…
Multi-View Correlation Distillation for Incremental Object Detection
Dongbao Yang, Yu Zhou, Weiping Wang
In real applications, new object classes often emerge after the detection model has been trained on a prepared dataset with fixed classes. Due to the storage burden and the privacy…
Context-Constrained Accurate Contour Extraction for Occlusion Edge Detection
Rui Lu, Menghan Zhou, Anlong Ming +1
Occlusion edge detection requires both accurate locations and context constraints of the contour. Existing CNN-based pipeline does not utilize adaptive methods to filter the noise…
PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
Ye Zhang, Yu Zhou, Jingwen Qi +11
Deep learning based automated pathological diagnosis has markedly improved diagnostic efficiency and reduced variability between observers, yet its clinical adoption remains limite…
Cyclic Delay-Doppler Shift: A Simple Transmit Diversity Technique for Ultra-Reliable Communications in Doubly Selective Channels
Haoran Yin, Yu Zhou, Yanqun Tang +7
Affine frequency division multiplexing (AFDM) and orthogonal time frequency space (OTFS) are two promising advanced waveforms proposed for reliable communications in high-mobility…
Observation of Anticorrelation with Classical Light in a Linear Optical System
Jianbin Liu, Yu Zhou, Fu-Li Li +1
Two-photon anticorrelation is observed when laser and pseudothermal light beams are incident to the two input ports of a Hong-Ou-Mandel interferometer, respectively. The spatial se…
Naked-Eye Ghost Imaging via Photoelectric-Feedback
Gao Wang, Huaibin Zheng, Yu Zhou +6
Based on optical correlations, ghost imaging is usually reconstructed by computer algorithm from the acquired data. We here proposed an alternatively high contrast naked-eye ghost…
MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
Xurui Li, Feng Xue, Yu Zhou
Zero-shot anomaly classification (AC) and segmentation (AS) methods aim to identify and outline defects without using any labeled samples. In this paper, we reveal a key property t…
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Yifei Zhang, Chang Liu, Jin Wei +4
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while t…
Study of the relativistic charged particle beam propagation in Earth's magnetic field
Meihua Fang, Zheng liang, Yingkui Gong +5
Relativistic charged particle beam can be used as destructive beam weapons in space for debris removal tasks. The trajectories of charged particles are affected by both electric an…
AIR tilting subcategories of extended hearts
Jiaqun Wei, Yu Zhou
We introduce the notion of AIR tilting subcategories of extended hearts of -structures on a triangulated category associated with silting subcategories. This notion generalizes…
Imaging around corners with single-pixel detector by computational ghost imaging
Bin Bai, Jianbin Liu, Yu Zhou +3
We have designed a single-pixel camera with imaging around corners based on computational ghost imaging. It can obtain the image of an object when the camera cannot look at the obj…
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
Yan Zhang, Gangyan Zeng, Huawen Shen +3
Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. I…
Study on Leveraging Wind Farm Reactive Power Potential for Uncertain Power System Reactive Power Optimization
Yu Zhou, Zhengshuo Li
This paper suggests leveraging reactive power potential (RPP) embedded in wind farms to improve power system operational safety and optimality. First, three typical RPP provision a…
Profitability of simple stationary technical trading rules with high-frequency data of Chinese Index Futures
Jing-Chao Chen, Yu Zhou, Xi Wang
Technical trading rules have been widely used by practitioners in financial markets for a long time. The profitability remains controversial and few consider the stationarity of te…
Security Analysis of Mode-Pairing Quantum Key Distribution with Flexible Pairing Strategy
Yi-Fei Lu, Yang Wang, Yan-Yang Zhou +10
Mode-pairing quantum key distribution (MP-QKD) is advantageous for long-distance secure communication, leveraging its simple implementation and quadratic scaling capacity. The post…
A Parallel Distributed Strategy for Arraying a Scattered Robot Swarm
Dominik Krupke, Michael Hemmer, James McLurkin +2
We consider the problem of organizing a scattered group of robots in two-dimensional space, with geometric maximum distance between robots. The communication graph of the s…
Fair Allocation of Indivisible Chores: Beyond Additive Costs
Bo Li, Fangxiao Wang, Yu Zhou
We study the maximin share (MMS) fair allocation of indivisible chores to agents who have costs for completing the assigned chores. It is known that exact MMS fairness cann…
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
Qilang Ye, Wei Zeng, Meng Liu +4
Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, whic…
Renewing Iterative Self-labeling Domain Adaptation with Application to the Spine Motion Prediction
Gecheng Chen, Yu Zhou, Xudong Zhang +1
The area of transfer learning comprises supervised machine learning methods that cope with the issue when the training and testing data have different input feature spaces or distr…
Second-order Fermionic Interference with Independent Photons
Jianbin Liu, Hui Chen, Yu Zhou +3
The experimental study of the second-order interference with fermions is much less than the one with bosons since it is much more difficult to do experiments with fermions than wit…
Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images
Yu Cai, Hao Chen, Xin Yang +2
Medical anomaly detection is a crucial yet challenging task aimed at recognizing abnormal images to assist in diagnosis. Due to the high-cost annotations of abnormal images, most m…
EviDAG: Auditable Causal DAG Authoring with Biomedical Literature
Yi-han Sheu, Michael R. Steigman, Yu Zhou +3
Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts must connect study variables to p…
MMF: Multi-Task Multi-Structure Fusion for Hierarchical Image Classification
Xiaoni Li, Yucan Zhou, Yu Zhou +1
Hierarchical classification is significant for complex tasks by providing multi-granular predictions and encouraging better mistakes. As the label structure decides its performance…
Even Order Explicit Symplectic Geometric Algorithms for Solving Quaternions in Guidance Navigation and Control via Diagonal Padé Approximation and Cayley Transform
Hong-Yan Zhang, Fei Liu, Yu Zhou +1
Quaternion kinematical differential equation (QKDE) plays a key role in navigation, control and guidance systems. Although explicit symplectic geometric algorithms (ESGA) for this…
Finite/fixed-time Stabilization of Linear Systems with States Quantization
Yu Zhou, Andrey Polyakov, Gang Zheng
This paper develops a homogeneity-based approach to finite/fixed-time stabilization of linear time-invariant (LTI) system with quantized measurements. A sufficient condition for fi…
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
Daiqing Wu, Dongbao Yang, Sicheng Zhao +2
Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabili…
DRIVE: Dockerfile Rule Mining and Violation Detection
Yu Zhou, Weilin Zhan, Zi Li +3
A Dockerfile defines a set of instructions to build Docker images, which can then be instantiated to support containerized applications. Recent studies have revealed a considerable…
Cluster combinatorics of d-cluster categories
Yu Zhou, Bin Zhu
We study the cluster combinatorics of cluster tilting objects in cluster categories. By using mutations of maximal rigid objects in cluster categories which are defined…
Superbunching pseudothermal light with intensity modulated laser light and rotating groundglass
Yu Zhou, Xuexing Zhang, Zhengpeng Wang +6
Pseudothermal light by scattering laser light from rotating groundglass has been extensively employed to study optical coherence in both classical and quantum optics ever since its…
Two-Level Residual Distillation based Triple Network for Incremental Object Detection
Dongbao Yang, Yu Zhou, Dayan Wu +3
Modern object detection methods based on convolutional neural network suffer from severe catastrophic forgetting in learning new classes without original data. Due to time consumpt…