Publications (40)
AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications
Daoce Wang, Jesus Pulido, Pascal Grosset +11
As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences expo…
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e…
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
Xinyi Li, Ang Li, Bo Fang +3
NVIDIA Tensor Cores and AMD Matrix Cores (together called Matrix Accelerators) are of growing interest in high-performance computing and machine learning owing to their high perfor…
Efficient Hierarchical State Vector Simulation of Quantum Circuits via Acyclic Graph Partitioning
Bo Fang, M. Yusuf Ãzkaya, Ang Li +2
Early but promising results in quantum computing have been enabled by the concurrent development of quantum algorithms, devices, and materials. Classical simulation of quantum prog…
Can Large Language Models Understand Intermediate Representations in Compilers?
Hailong Jiang, Jianfeng Zhu, Yao Wan +4
Intermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by Large Language Models (LLMs) remains underexplored. In t…
Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?
Wenhao Wu, Haipeng Luo, Bo Fang +2
Most existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences. However, in real-world scenarios, online…
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
Yuhang Liang, Xinyi Li, Jie Ren +3
Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensi…
ScSAM: Debiasing Morphology and Distributional Variability in Subcellular Semantic Segmentation
Bo Fang, Jianan Fan, Dongnan Liu +4
The significant morphological and distributional variability among subcellular components poses a long-standing challenge for learning-based organelle segmentation models, signific…
A Hybrid System for Learning Classical Data in Quantum States
Samuel A. Stein, Ryan L'Abbate, Wenrui Mu +6
Deep neural network powered artificial intelligence has rapidly changed our daily life with various applications. However, as one of the essential steps of deep neural networks, tr…
VeriHGN: Heterogeneous Graph-Based Congestion Prediction for Chip Layout Verification
Runbang Hu, Bo Fang, Bingzhe Li +1
As Very Large Scale Integration (VLSI) designs continue to scale in size and complexity, layout verification has become a central challenge in modern Electronic Design Automation (…
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
Haoyuan Sun, Jing Wang, Yuxin Song +9
Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for furth…
TensorFI: A Flexible Fault Injection Framework for TensorFlow Applications
Zitao Chen, Niranjhana Narayanan, Bo Fang +3
As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prio…
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Xinyao Zhang, Wenkai Dong, Yuxin Song +10
Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely…
Characterizing Impacts of Storage Faults on HPC Applications: A Methodology and Insights
Bo Fang, Daoce Wang, Sian Jin +6
In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses,…
Video 3D Sampling for Self-supervised Representation Learning
Wei Li, Dezhao Luo, Bo Fang +2
Most of the existing video self-supervised methods mainly leverage temporal signals of videos, ignoring that the semantics of moving objects and environmental information are all c…
Final Report for CHESS: Cloud, High-Performance Computing, and Edge for Science and Security
Nathan Tallent, Jan Strube, Luanzheng Guo +9
Automating the theory-experiment cycle requires effective distributed workflows that utilize a computing continuum spanning lab instruments, edge sensors, computing resources at mu…
MEMQSim: Highly Memory-Efficient and Modularized Quantum State-Vector Simulation
Boyuan Zhang, Bo Fang, Qiang Guan +2
In this extended abstract, we have introduced a highly memory-efficient state vector simulation of quantum circuits premised on data compression, harnessing the capabilities of bot…
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
Boyuan Zhang, Bo Fang, Fanjiang Ye +4
Full-state quantum circuit simulation requires exponentially increased memory size to store the state vector as the number of qubits scales, presenting significant limitations in c…
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu +3
Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning deri…
DistinctAD: Distinctive Audio Description Generation in Contexts
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automa…
Toward Lossless Homomorphic Encryption for Scientific Computation
Muhammad Jahanzeb Khan, Bo Fang, Dongfang Zhao
This paper presents a comprehensive investigation into encrypted computations using the CKKS (Cheon-Kim-Kim-Song) scheme, with a focus on multi-dimensional vector operations and re…
Broadband Cross-Circular Polarization Carpet Cloaking based on a Phase Change Material Metasurface in the Mid-infrared Region
Bo Fang, Dantian Feng, Peng Chen +6
In view of the fact that most invisibility devices focus on linear polarization cloaking and that the characteristics of mid infrared cloaking are rarely studied, we propose a cros…
Sensing performance enhancement via asymmetric gain optimization in the atom-light hybrid interferometer
Zhifei Yu, Bo Fang, Shuying Chen +4
The SU (1,1)-type atom-light hybrid interferometer (SALHI) is a kind of interferometer that is sensitive to both the optical phase and atomic phase. However, the loss has been an u…
Quantum Sampling Architecture for Protein Structure Reconstruction on Utility-Scale Hardware
Yuqi Zhang, Bo Fang, Yuxin Yang +6
Predicting the structure of short peptides in protein binding pockets remains difficult because this regime requires physics-based conformational search, yet existing methods do no…
UATVR: Uncertainty-Adaptive Text-Video Retrieval
Bo Fang, Wenhao Wu, Chang Liu +6
With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted…
Understanding the Landscape of Ampere GPU Memory Errors
Zhu Zhu, Yu Sun, Dhatri Parakal +9
Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an ess…
Exploring Relations in Untrimmed Videos for Self-Supervised Learning
Dezhao Luo, Bo Fang, Yu Zhou +3
Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this se…
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
Zachary Coalson, Jeonghyun Woo, Chris S. Lin +8
We study a new vulnerability in commercial-scale safety-aligned large language models (LLMs): their refusal to generate harmful responses can be broken by flipping only a few bits…
A Survey on Agentic Multimodal Large Language Models
Huanjin Yao, Ruifei Zhang, Jiaxing Huang +8
With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI…
MARS: Malleable Actor-Critic Reinforcement Learning Scheduler
Betis Baheri, Jacob Tronge, Bo Fang +3
In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermed…
QDockBank: A Dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers
Yuqi Zhang, Yuxin Yang, Cheng-Chang Lu +4
Protein structure prediction is a core challenge in computational biology, particularly for fragments within ligand-binding regions, where accurate modeling is still difficult. Qua…
Graph-VQE: A CUDA-Q Multi-QPU Simulation Framework for Hamiltonian-Aware Protein-Folding VQE
Yujun Feng, Yuqi Zhang, Jingyi Huang +4
The Variational Quantum Eigensolver (VQE) is essential for molecular simulation in drug discovery, but hardware noise and algorithmic limits restrict its precision. While the NVIDI…
Splaxel: Efficient Distributed Training of 3D Gaussian Splatting for Large-scale Scene Reconstruction via Pixel-level Communication
Wenqi Jia, Zhewen Hu, Ying Huang +10
3D Gaussian Splatting (3DGS) enables high-fidelity and real-time 3D scene reconstruction, but scaling training to large-scale scenes requires optimizing hundreds of millions of Gau…
MPGemmFI: A Fault Injection Technique for Mixed Precision GEMM in ML Applications
Bo Fang, Xinyi Li, Harvey Dam +9
Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). To meet such demand, one of the critical features of machine-learning-specific accelerator…
M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking
Qiangqiang Wu, Tianyu Yang, Bo Fang +4
Tracking Any Point (TAP) has emerged as a fundamental tool for video understanding. Current approaches adapt Vision Foundation Models (VFMs) like DINOv2 via offline finetuning or t…
TANQ-Sim: Tensorcore Accelerated Noisy Quantum System Simulation via QIR on Perlmutter HPC
Ang Li, Chenxu Liu, Samuel Stein +7
Although there have been remarkable advances in quantum computing (QC), it remains crucial to simulate quantum programs using classical large-scale parallel computing systems to va…
Red-QAOA: Efficient Variational Optimization through Circuit Reduction
Meng Wang, Bo Fang, Ang Li +1
The Quantum Approximate Optimization Algorithm (QAOA) addresses combinatorial optimization challenges by converting inputs to graphs. However, the optimal parameter searching proce…
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Bo Fang, Xinyao Zhang, Yuxin Song +3
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
Zachary Coalson, Bo Fang, Sanghyun Hong
Multi-turn interaction length is a dominant factor in the operational costs of conversational LLMs. In this work, we present a new failure mode in conversational LLMs: turn amplifi…
QuGAN: A Quantum State Fidelity based Generative Adversarial Network
Samuel A. Stein, Betis Baheri, Daniel Chen +5
Tremendous progress has been witnessed in artificial intelligence where neural network backed deep learning systems have been used, with applications in almost every domain. As a r…