papers

Publications (40)

cs.DC2023

AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications

Daoce Wang, Jesus Pulido, Pascal Grosset +11

As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences expo…

cs.CV2025

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Bo Fang, Wenhao Wu, Qiangqiang Wu +2

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e…

cs.AR2024

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

Xinyi Li, Ang Li, Bo Fang +3

NVIDIA Tensor Cores and AMD Matrix Cores (together called Matrix Accelerators) are of growing interest in high-performance computing and machine learning owing to their high perfor…

quant-ph2022

Efficient Hierarchical State Vector Simulation of Quantum Circuits via Acyclic Graph Partitioning

Bo Fang, M. Yusuf Özkaya, Ang Li +2

Early but promising results in quantum computing have been enabled by the concurrent development of quantum algorithms, devices, and materials. Classical simulation of quantum prog…

cs.LG2025

Can Large Language Models Understand Intermediate Representations in Compilers?

Hailong Jiang, Jianfeng Zhu, Yao Wan +4

Intermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by Large Language Models (LLMs) remains underexplored. In t…

cs.CV2023

Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?

Wenhao Wu, Haipeng Luo, Bo Fang +2

Most existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences. However, in real-world scenarios, online…

cs.DC2025

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Yuhang Liang, Xinyi Li, Jie Ren +3

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensi…

cs.CV2025

ScSAM: Debiasing Morphology and Distributional Variability in Subcellular Semantic Segmentation

Bo Fang, Jianan Fan, Dongnan Liu +4

The significant morphological and distributional variability among subcellular components poses a long-standing challenge for learning-based organelle segmentation models, signific…

quant-ph2021

A Hybrid System for Learning Classical Data in Quantum States

Samuel A. Stein, Ryan L'Abbate, Wenrui Mu +6

Deep neural network powered artificial intelligence has rapidly changed our daily life with various applications. However, as one of the essential steps of deep neural networks, tr…

cs.AR2026

VeriHGN: Heterogeneous Graph-Based Congestion Prediction for Chip Layout Verification

Runbang Hu, Bo Fang, Bingzhe Li +1

As Very Large Scale Integration (VLSI) designs continue to scale in size and complexity, layout verification has become a central challenge in modern Electronic Design Automation (…

cs.CV2026

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Haoyuan Sun, Jing Wang, Yuxin Song +9

Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for furth…

cs.DC2020

TensorFI: A Flexible Fault Injection Framework for TensorFlow Applications

Zitao Chen, Niranjhana Narayanan, Bo Fang +3

As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prio…

cs.CV2026

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

Xinyao Zhang, Wenkai Dong, Yuxin Song +10

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely…

cs.DC2021

Characterizing Impacts of Storage Faults on HPC Applications: A Methodology and Insights

Bo Fang, Daoce Wang, Sian Jin +6

In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses,…

cs.CV2021

Video 3D Sampling for Self-supervised Representation Learning

Wei Li, Dezhao Luo, Bo Fang +2

Most of the existing video self-supervised methods mainly leverage temporal signals of videos, ignoring that the semantics of moving objects and environmental information are all c…

cs.DC2024

Final Report for CHESS: Cloud, High-Performance Computing, and Edge for Science and Security

Nathan Tallent, Jan Strube, Luanzheng Guo +9

Automating the theory-experiment cycle requires effective distributed workflows that utilize a computing continuum spanning lab instruments, edge sensors, computing resources at mu…

quant-ph2023

MEMQSim: Highly Memory-Efficient and Modularized Quantum State-Vector Simulation

Boyuan Zhang, Bo Fang, Qiang Guan +2

In this extended abstract, we have introduced a highly memory-efficient state vector simulation of quantum circuits premised on data compression, harnessing the capabilities of bot…

cs.DC2024

Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework

Boyuan Zhang, Bo Fang, Fanjiang Ye +4

Full-state quantum circuit simulation requires exponentially increased memory size to store the state vector as the number of qubits scales, presenting significant limitations in c…

cs.CV2025

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Bo Fang, Yuxin Song, Qiangqiang Wu +3

Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning deri…

cs.CV2024

DistinctAD: Distinctive Audio Description Generation in Contexts

Bo Fang, Wenhao Wu, Qiangqiang Wu +2

Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automa…

cs.CR2023

Toward Lossless Homomorphic Encryption for Scientific Computation

Muhammad Jahanzeb Khan, Bo Fang, Dongfang Zhao

This paper presents a comprehensive investigation into encrypted computations using the CKKS (Cheon-Kim-Kim-Song) scheme, with a focus on multi-dimensional vector operations and re…

physics.optics2022

Broadband Cross-Circular Polarization Carpet Cloaking based on a Phase Change Material Metasurface in the Mid-infrared Region

Bo Fang, Dantian Feng, Peng Chen +6

In view of the fact that most invisibility devices focus on linear polarization cloaking and that the characteristics of mid infrared cloaking are rarely studied, we propose a cros…

quant-ph2022

Sensing performance enhancement via asymmetric gain optimization in the atom-light hybrid interferometer

Zhifei Yu, Bo Fang, Shuying Chen +4

The SU (1,1)-type atom-light hybrid interferometer (SALHI) is a kind of interferometer that is sensitive to both the optical phase and atomic phase. However, the loss has been an u…

cs.ET2026

Quantum Sampling Architecture for Protein Structure Reconstruction on Utility-Scale Hardware

Yuqi Zhang, Bo Fang, Yuxin Yang +6

Predicting the structure of short peptides in protein binding pockets remains difficult because this regime requires physics-based conformational search, yet existing methods do no…

cs.CV2023

UATVR: Uncertainty-Adaptive Text-Video Retrieval

Bo Fang, Wenhao Wu, Chang Liu +6

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted…

cs.DC2025

Understanding the Landscape of Ampere GPU Memory Errors

Zhu Zhu, Yu Sun, Dhatri Parakal +9

Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an ess…

cs.CV2020

Exploring Relations in Untrimmed Videos for Self-Supervised Learning

Dezhao Luo, Bo Fang, Yu Zhou +3

Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this se…

cs.CR2025

PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips

Zachary Coalson, Jeonghyun Woo, Chris S. Lin +8

We study a new vulnerability in commercial-scale safety-aligned large language models (LLMs): their refusal to generate harmful responses can be broken by flipping only a few bits…

cs.CV2025

A Survey on Agentic Multimodal Large Language Models

Huanjin Yao, Ruifei Zhang, Jiaxing Huang +8

With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI…

cs.DC2022

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Betis Baheri, Jacob Tronge, Bo Fang +3

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermed…

cs.ET2025

QDockBank: A Dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers

Yuqi Zhang, Yuxin Yang, Cheng-Chang Lu +4

Protein structure prediction is a core challenge in computational biology, particularly for fragments within ligand-binding regions, where accurate modeling is still difficult. Qua…

cs.ET2026

Graph-VQE: A CUDA-Q Multi-QPU Simulation Framework for Hamiltonian-Aware Protein-Folding VQE

Yujun Feng, Yuqi Zhang, Jingyi Huang +4

The Variational Quantum Eigensolver (VQE) is essential for molecular simulation in drug discovery, but hardware noise and algorithmic limits restrict its precision. While the NVIDI…

cs.DC2026

Splaxel: Efficient Distributed Training of 3D Gaussian Splatting for Large-scale Scene Reconstruction via Pixel-level Communication

Wenqi Jia, Zhewen Hu, Ying Huang +10

3D Gaussian Splatting (3DGS) enables high-fidelity and real-time 3D scene reconstruction, but scaling training to large-scale scenes requires optimizing hundreds of millions of Gau…

cs.DC2023

MPGemmFI: A Fault Injection Technique for Mixed Precision GEMM in ML Applications

Bo Fang, Xinyi Li, Harvey Dam +9

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). To meet such demand, one of the critical features of machine-learning-specific accelerator…

cs.CV2026

M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking

Qiangqiang Wu, Tianyu Yang, Bo Fang +4

Tracking Any Point (TAP) has emerged as a fundamental tool for video understanding. Current approaches adapt Vision Foundation Models (VFMs) like DINOv2 via offline finetuning or t…

quant-ph2024

TANQ-Sim: Tensorcore Accelerated Noisy Quantum System Simulation via QIR on Perlmutter HPC

Ang Li, Chenxu Liu, Samuel Stein +7

Although there have been remarkable advances in quantum computing (QC), it remains crucial to simulate quantum programs using classical large-scale parallel computing systems to va…

quant-ph2024

Red-QAOA: Efficient Variational Optimization through Circuit Reduction

Meng Wang, Bo Fang, Ang Li +1

The Quantum Approximate Optimization Algorithm (QAOA) addresses combinatorial optimization challenges by converting inputs to graphs. However, the optimal parameter searching proce…

cs.CV2026

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

Bo Fang, Xinyao Zhang, Yuxin Song +3

Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…

cs.LG2026

Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs

Zachary Coalson, Bo Fang, Sanghyun Hong

Multi-turn interaction length is a dominant factor in the operational costs of conversational LLMs. In this work, we present a new failure mode in conversational LLMs: turn amplifi…

quant-ph2022

QuGAN: A Quantum State Fidelity based Generative Adversarial Network

Samuel A. Stein, Betis Baheri, Daniel Chen +5

Tremendous progress has been witnessed in artificial intelligence where neural network backed deep learning systems have been used, with applications in almost every domain. As a r…