Publications (94)
Rethinking Code Performance Benchmarks for LLMs
Nhat Minh Le, Yisen Xu, Zhijie Wang +2
Many function-level performance benchmarks have been proposed to evaluate whether large language models (LLMs) can generate efficient programs. However, results on these benchmarks…
Screencast-Based Analysis of User-Perceived GUI Responsiveness
Wei Liu, Linqiang Guo, Yi Wen Heng +4
GUI responsiveness is critical for a positive user experience in mobile applications. Even brief delays in visual feedback can frustrate users and lead to negative reviews. However…
PrivScope: Task-scoped Disclosure Control for Hybrid Agentic Systems
Shafizur Rahman Seeam, Zhengxiong Li, Zhiyuan Yu +3
Hybrid local--cloud agents enrich user requests with context from persistent working state before delegating capability-intensive subtasks to a cloud language model (CLM). While th…
Bug Report Specification Refinement with Trajectory Guidance for Automated Program Repair
S M Farah Al Fahim, Md Nakhla Rafi, Md Ahasanuzzaman +5
Bug reports serve as task specifications for repository-level automated program repair (APR) agents, but they often describe only the observed failure and omit repair-relevant info…
Turning Interaction History into Execution State: A Runtime Layer for Long-Horizon Coding Agents
Zehao Wang, Yisen Xu, Chenglin Li +5
Long-horizon coding agents accumulate hundreds of actions and observations in their trajectories, yet nothing in this record indicates which observations still describe the reposit…
A Stable FDTD Subgridding Scheme with SBP-SAT for Transient Electromagnetic Analysis
Yu Cheng, Yuhui Wang, Hanhong Liu +5
We proposed a provably stable FDTD subgridding method for accurate and efficient transient electromagnetic analysis. In the proposed method, several field components are properly a…
The Causal Impact of Dean's List Recognition on Academic Performance: Evidence from a Regression Discontinuity Design
Luc, Chen
This study examines the causal impact of being placed on the Dean's List, a positive education incentive, on future student performance using a regression discontinuity design. The…
Development and Validation of a Deep Learning Algorithm for Improving Gleason Scoring of Prostate Cancer
Kunal Nagpal, Davis Foote, Yun Liu +17
For prostate cancer patients, the Gleason score is one of the most important prognostic factors, potentially determining treatment independent of the stage. However, Gleason scorin…
Dialogue Act Patterns in GenAI-Mediated L2 Oral Practice: A Sequential Analysis of Learner-Chatbot Interactions
Liqun He, Shijun, Chen +2
While generative AI (GenAI) voice chatbots offer scalable opportunities for second language (L2) oral practice, the interactional processes related to learners' gains remain undere…
SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents
Feng Lin, Dong Jae Kim, Tse-Husn +1
Software process models are essential to facilitate collaboration and communication among software teams to solve complex development tasks. Inspired by these software engineering…
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
Chia-Pei, Chen, Kentaroh Toyoda +2
Web-browsing AI agents are increasingly deployed in enterprise settings under strict whitelists of approved domains, yet adversaries can still influence them by embedding hidden in…
Asymmetric price adjustment over the business cycle
Daniel Levy, Haipeng, Chen +6
Studies of micro-level price datasets find more frequent small price increases than decreases, which can be explained by consumer inattention because time-constrained shoppers migh…
Diffusion Large Language Models for Black-Box Optimization
Ye Yuan, Can, Chen +4
Offline black-box optimization (BBO) aims to find optimal designs based solely on an offline dataset of designs and their labels. Such scenarios frequently arise in domains like DN…
Multimodal Latent Fusion of ECG Leads for Early Assessment of Pulmonary Hypertension
Mohammod N. I. Suvon, Shuo Zhou, Prasun C. Tripathi +7
Recent advancements in early assessment of pulmonary hypertension (PH) primarily focus on applying machine learning methods to centralized diagnostic modalities, such as 12-lead el…
Feature selection algorithm based on incremental mutual information and cockroach swarm optimization
Zhao, Chen
Feature selection is an effective preprocessing technique to reduce data dimension. For feature selection, rough set theory provides many measures, among which mutual information i…
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
Rabeya Khatun Muna, Md Nakhla Rafi, Tse-Hsun +1
Continuous Integration (CI) enforces repository-level correctness through multi-stage workflows and is central to modern software development, yet diagnosing and repairing CI failu…
AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding
Shuang Liang, Hao Mark Chen, Hao +6
Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verification can activate the union…
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
Chung-En, Yu, Hsuan-Chih +3
To develop trustworthy Vision-Language Models (VLMs), it is essential to address adversarial robustness and hallucination mitigation, both of which impact factual accuracy in high-…
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
Chenglin Li, Yisen Xu, Zehao Wang +3
Repository-level automated program repair (APR) requires long-horizon reasoning over interdependent decisions. However, most LLM-based approaches reconstruct repair reasoning indep…
Temporal Relevance Analysis for Video Action Models
Quanfu Fan, Donghyun Kim, Chun-Fu +4
In this paper, we provide a deep analysis of temporal modeling for action recognition, an important but underexplored problem in the literature. We first propose a new approach to…
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
Yisen Xu, Jinqiu Yang, Tse-Hsun +1
Large Language Models (LLMs) have recently attracted wide interest for tackling software engineering tasks. In contrast to code generation, refactoring demands precise, semantics-p…
MobileUPReg: Identifying User-Perceived Performance Regressions in Mobile OS Versions
Wei Liu, Yi Wen Heng, Feng Lin +3
Mobile operating systems (OS) are frequently updated, but such updates can unintentionally degrade user experience by introducing performance regressions. Existing detection techni…
Agent-SAMA: State-Aware Mobile Assistant
Linqiang Guo, Wei Liu, Yi Wen Heng +3
Mobile Graphical User Interface (GUI) agents aim to autonomously complete tasks within or across apps based on user instructions. While recent Multimodal Large Language Models (MLL…
Synergizing Self-Regulation and Artificial-Intelligence Literacy Towards Future Human-AI Integrative Learning
Long, Zhang, Shijun +1
Self-regulated learning (SRL) and Artificial-Intelligence (AI) literacy are becoming key competencies for successful human-AI interactive learning, vital to future education. Howev…
Physics-Informed Neural Operator for Electromagnetic Inverse Scattering Problems
Q. C. Dong, Zi-Xuan Su, Qing Huo Liu +3
This paper proposes a physics-informed neural operator (PINO) framework for solving inverse scattering problems, enabling rapid and accurate reconstructions under diverse measureme…
Entropy-based Thermal Sensor Placement and Temperature Reconstruction based on Adaptive Compressive Sensing Theory
Kun-Chih, Chen, Chia-Hsin Chen +2
This paper addresses the challenges of thermal sensor allocation and full-chip temperature reconstruction in multi-core systems by leveraging an entropy-based sensor placement stra…
Future Physics Programme of BESIII
M. Ablikim, M. N. Achasov, P. Adlarson +482
There has recently been a dramatic renewal of interest in the subjects of hadron spectroscopy and charm physics. This renaissance has been driven in part by the discovery of a plet…
Generalizing matrix representations to fully heterochronous ranked tree shapes
Chris Jennings-Shaffer, Ziyue, Chen +2
Phylogenetic tree shapes capture fundamental signatures of evolution. We consider ``ranked'' tree shapes, which are equipped with a total order on the internal nodes compatible wit…
GraphCompNet: A Position-Aware Model for Predicting and Compensating Shape Deviations in 3D Printing
Juheon Lee, Lei, Chen +3
Shape deviation modeling and compensation in additive manufacturing are pivotal for achieving high geometric accuracy and enabling industrial-scale production. Critical challenges…
A K-fold Method for Baseline Estimation in Policy Gradient Algorithms
Nithyanand Kota, Abhishek Mishra, Sunil Srinivasa +3
The high variance issue in unbiased policy-gradient methods such as VPG and REINFORCE is typically mitigated by adding a baseline. However, the baseline fitting itself suffers from…
Retrieval-Oriented Code Representations in Agentic Bug Localization
Genevieve Caumartin, Tse-Hsun, Chen +1
The paper evaluates how different code representations, including LLM-generated textual summaries, affect the effectiveness and cost of file-level bug localization, finding that ro…
MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
Yisen Xu, Feng Lin, Jinqiu Yang +3
Maintaining and scaling software systems relies heavily on effective code refactoring, yet this process remains labor-intensive, requiring developers to carefully analyze existing…
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
Renyuan Liu, Yuyang Leng, Kaiyan Liu +6
Recent advancements in on-device training for deep neural networks have underscored the critical need for efficient activation compression to overcome the memory constraints of mob…
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
Linqiang Guo, Wei Liu, Li Gu +3
Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action,…
PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing
Linqiang Guo, Wei Liu, Yi Wen Heng +3
Graphical User Interfaces (GUIs) are the primary means by which users interact with mobile applications, making them crucial to both app functionality and user experience. However,…
A Hybrid SIE-PDE Formulation Without Boundary Condition Requirement for Transverse Magnetic Electromagnetic Analysis
Aipeng Sun, Zekun Zhu, Shunchuan Yang +2
A hybrid surface integral equation partial differential equation (SIE-PDE) formulation without the boundary condition requirement is proposed to solve the transverse magnetic (TM)…
Short-Term Forecasting of Passenger Demand under On-Demand Ride Services: A Spatio-Temporal Deep Learning Approach
Jintao Ke, Hongyu Zheng, Hai Yang +2
Short-term passenger demand forecasting is of great importance to the on-demand ride service platform, which can incentivize vacant cars moving from over-supply regions to over-dem…
A Survey of Intrusion Detection Systems Leveraging Host Data
Tarrah R. Glass-Vanderlan, Michael D. Iannacone, Maria S. Vincent +3
This survey focuses on intrusion detection systems (IDS) that leverage host-based data sources for detecting attacks on enterprise network. The host-based IDS (HIDS) literature is…
Physics-Informed Geometric Operators to Support Surrogate, Dimension Reduction and Generative Models for Engineering Design
Shahroz Khan, Zahid Masood, Muhammad Usama +4
In this work, we propose a set of physics-informed geometric operators (GOs) to enrich the geometric data provided for training surrogate/discriminative models, dimension reduction…
Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
Nolan Dey, Gurpreet Gosal, Zhiming +6
We study recent research advances that improve large language models through efficient pre-training and scaling, and open datasets and tools. We combine these advances to introduce…
Measuring similarity between two mixture trees using mixture distance metric and algorithms
Justie Su-Tzu Juan, Yi-Ching Chen, Chen-Hui Lin +2
Ancestral mixture model, proposed by Chen and Lindsay (2006), is an important model to build a hierarchical tree from high dimensional binary sequences. Mixture trees created from…
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
Feng Lin, Dong Jae Kim, Zhenhao Li +3
When using LLMs to address Non-Functional Requirements (NFRs), developers may behave differently (e.g., expressing the same NFR in different words). Robust LLMs should output consi…
Slicing Vision Transformer for Flexible Inference
Yitian Zhang, Huseyin Coskun, Xu Ma +6
Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe…
Missing-Modality-Aware Graph Neural Network for Cancer Classification
Sina Tabakhi, Chen, Haiping Lu
A key challenge in learning from multimodal biological data is missing modalities, where data from one or more modalities are absent for some patients. Existing approaches either e…
Average-case Speedup for Product Formulas
Chi-Fang, Chen, Fernando G. S. L. Brandão
Quantum simulation is a promising application of future quantum computers. Product formulas, or Trotterization, are the oldest and still remain an appealing method to simulate quan…
An Efficient Outlier Detection Algorithm for Data Streaming
Rui Hu, Luc, Chen +1
The nature of modern data is increasingly real-time, making outlier detection crucial in any data-related field, such as finance for fraud detection and healthcare for monitoring p…
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
Chengzhi Shen, Weixiang Shen, Tobias Susetzky +7
Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, undersco…
Dynamic Matching Under Patience Imbalance
Zhiyuan Chen, Rui, Chen +2
We study a dynamic matching problem on a two-sided platform with unbalanced patience, in which long-lived supply accumulates over time with a unit waiting cost per period, while sh…
PASS: Private Attributes Protection with Stochastic Data Substitution
Yizhuo Chen, Chun-Fu, Chen +3
The growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people's private information irrelevant to the services. Vari…
Studying and Recommending Information Highlighting in Stack Overflow Answers
Shahla Shaan Ahmed, Shaowei Wang, Yuan Tian +3
Context: Navigating the knowledge of Stack Overflow (SO) remains challenging. To make the posts vivid to users, SO allows users to write and edit posts with Markdown or HTML so tha…
Can An Image Classifier Suffice For Action Recognition?
Quanfu Fan, Chun-Fu, Chen +1
We explore a new perspective on video understanding by casting the video recognition problem as an image recognition task. Our approach rearranges input video frames into super ima…
Radio-Based Passive Target Tracking by a Mobile Receiver with Unknown Transmitter Position
Ke Xu, Rui Zhang, He +1
In this paper, we propose a radio-based passive target tracking algorithm using multipath measurements, including the angle of arrival and relative distance. We focus on a scenario…
Zero-Ending Prices, Cognitive Convenience, and Price Rigidity
Avichai Snir, Haipeng, Chen +1
We assess the role of cognitive convenience in the popularity and rigidity of 0 ending prices in convenience settings. Studies show that 0 ending prices are common at convenience s…
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
Arjun Nichani, Hsiang Hsu, Chun-Fu +2
Fairness and privacy are two vital pillars of trustworthy machine learning. Despite extensive research on these individual topics, their relationship has received significantly les…
MSTGD:A Memory Stochastic sTratified Gradient Descent Method with an Exponential Convergence Rate
Aixiang, Chen, Jinting Zhang +2
The fluctuation effect of gradient expectation and variance caused by parameter update between consecutive iterations is neglected or confusing by current mainstream gradient optim…
An Unconditionally Stable Conformal LOD-FDTD Method For Curved PEC Objects and Its Application to EMC Problems
Hanhong Liu, Xiaoying Zhao, Xiang-Hua Wang +3
The traditional finite-difference time-domain (FDTD) method is constrained by the Courant-Friedrich-Levy (CFL) condition and suffers from the notorious staircase error in electroma…
Spatial-Temporal Inference of Urban Traffic Emissions Based on Taxi Trajectories and Multi-Source Urban Data
Jielun Liu, Ke Han, Xiqun +2
Vehicle trajectory data collected via GPS-enabled devices have played increasingly important roles in estimating network-wide traffic, given their broad spatial-temporal coverage a…
Neuron's Eye View: Inferring Features of Complex Stimuli from Neural Responses
Xin, Chen, Jeffrey M Beck +1
Experiments that study neural encoding of stimuli at the level of individual neurons typically choose a small set of features present in the world --- contrast and luminance for vi…
Inverse Design of Nonlinear Mechanics of Bio-inspired Materials Through Interface Engineering and Bayesian Optimization
Wei Zhang, Mingjian Tang, Haoxuan Mu +6
In many biological materials such as nacre and bone, the material structure consists of hard grains and soft interfaces, with the interfaces playing a significant role in the mater…
Discovery of Timeline and Crowd Reaction of Software Vulnerability Disclosures
Yi Wen Heng, Zeyang Ma, Haoxiang Zhang +3
Reusing third-party libraries increases productivity and saves time and costs for developers. However, the downside is the presence of vulnerabilities in those libraries, which can…
TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance
Kan Wu, Houwen Peng, Zhenghong Zhou +10
In this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: af…
Vector Single-Source Surface Integral Equation for TE Scattering From Cylindrical Multilayered Objects
Zekun Zhu, Xiaochao Zhou, Shunchuan Yang +2
A single-source surface integral equation (SS-SIE) for transverse electric (TE) scattering from cylindrical multilayered objects is proposed in this paper. By incorporating the dif…
Seeing Beyond the Image: ECG and Anatomical Knowledge-Guided Myocardial Scar Segmentation from Late Gadolinium-Enhanced Images
Farheen Ramzan, Yusuf Kiberu, Nikesh Jathanna +5
Accurate segmentation of myocardial scar from late gadolinium enhanced (LGE) cardiac MRI is essential for evaluating tissue viability, yet remains challenging due to variable contr…
End-to-End Fairness Optimization with Fair Decision-Focused Learning
Yu Wang, Violet, Chen
Many real-world systems rely on predictive models to inform decisions, and fairness concerns arise in both the prediction and decision stages. We introduce end-to-end fairness opti…
Virtual Foundry Graphnet for Metal Sintering Deformation Prediction
Rachel, Chen, Juheon Lee +4
Metal Sintering is a necessary step for Metal Injection Molded parts and binder jet such as HP's metal 3D printer. The metal sintering process introduces large deformation varying…
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
Tse-Hsun, Chen
Software Engineering (SE) agents have shown promising abilities in supporting various SE tasks. Current SE agents remain fundamentally reactive, making decisions mainly based on co…
3D object quality prediction for Metal Jet Printer with Multimodal thermal encoder
Rachel, Chen, Wenjia Zheng +3
With the advancements in 3D printing technologies, it is extremely important that the quality of 3D printed objects, and dimensional accuracies should meet the customer's specifica…
Integrating Large Language Models in Financial Investments and Market Analysis: A Survey
Sedigheh Mahdavi, Jiating, Chen +3
Large Language Models (LLMs) have been employed in financial decision making, enhancing analytical capabilities for investment strategies. Traditional investment strategies often u…
Deployment-Time Memorization in Foundation-Model Agents
Lei, Chen, Guilin Zhang +8
Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a p…
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
Xingzhong Zhao, Ziqian Xie, Islam +5
Motivation: Modern bioinformatics workflows, particularly in imaging and representation learning, can generate thousands to tens of thousands of quantitative phenotypes from a sing…
Single-Source SIE for Two-Dimensional Arbitrarily Connected Penetrable and PEC Objects with Nonconformal Meshes
Zekun Zhu, Aipeng Sun, Xiaochao Zhou +3
We proposed a simple and efficient modular single-source surface integral equation (SS-SIE) formulation for electromagnetic analysis of arbitrarily connected penetrable and perfect…
An Empirical Study of Obsolete Answers on Stack Overflow
Haoxiang Zhang, Shaowei Wang, Tse-Hsun +3
Stack Overflow accumulates an enormous amount of software engineering knowledge. However, as time passes, certain knowledge in answers may become obsolete. Such obsolete answers, i…
Establishing Secrecy Region for Directional Modulation Scheme with Random Frequency Diverse Array
Shengping Lv, Jinsong Hu, Youjia Chen +3
Random frequency diverse array (RFDA) based directional modulation (DM) was proposed as a promising technology in secure communications to achieve a precise transmission of confide…
Towards Better Semantic Understanding of Mobile Interfaces
Srinivas Sunkara, Maria Wang, Lijuan Liu +6
Improving the accessibility and automation capabilities of mobile devices can have a significant positive impact on the daily lives of countless users. To stimulate research in thi…
SBEST: Spectrum-Based Fault Localization Without Fault-Triggering Tests
Md Nakhla Rafi, Lorena Barreto Simedo Pacheco, An Ran Chen +3
Fault localization is a critical step in software maintenance. Yet, many existing techniques, such as Spectrum-Based Fault Localization (SBFL), rely heavily on the availability of…
Dynamic Pricing in a Dual Market Environment
Wen, Chen, Adam Fleischhacker +1
This paper is concerned with the determination of pricing strategies for a firm that in each period of a finite horizon receives replenishment quantities of a single product which…
Software Design Document, Testing, Deployment and Configuration Management of the UUIS--a Team 2 COMP5541-W10 Project Approach
Omer Shahid Ahmad, Faisal Alrashdi, Jason +6
The Software Design Document of UUIS describes the prototype design details of the system architecture, database layer, deployment and configuration details as well as test cases p…
Whole-Slide Image Focus Quality: Automatic Assessment and Impact on AI Cancer Detection
Timo Kohlberger, Yun Liu, Melissa Moran +6
Digital pathology enables remote access or consults and powerful image analysis algorithms. However, the slide digitization process can create artifacts such as out-of-focus (OOF).…
Bilingual Adaptation of Monolingual Foundation Models
Gurpreet Gosal, Yishi Xu, Gokul Ramakrishnan +19
We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations…
Probe to Generate: Program Variant-Guided Test Augmentation for Repository-Level Repair Benchmarks
Chenglin Li, Yisen Xu, Zehao Wang +3
Test-based benchmarks such as SWE-bench have become a standard basis for evaluating automated issue resolution agents, deeming a patch correct if it passes a provided regression te…
Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity Estimation
Hsiang Hsu, Guihong Li, Shaohan Hu +2
Predictive multiplicity refers to the phenomenon in which classification tasks may admit multiple competing models that achieve almost-equally-optimal performance, yet generate con…
RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation
Zhen Li, Guenevere, Chen +3
Source code authorship attribution is an important problem often encountered in applications such as software forensics, bug fixing, and software quality analysis. Recent studies s…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
FogROS: An Adaptive Framework for Automating Fog Robotics Deployment
Kaiyuan, Chen, Yafei Liang +6
As many robot automation applications increasingly rely on multi-core processing or deep-learning models, cloud computing is becoming an attractive and economically viable resource…
Studying the Impact of Early Test Termination Due to Assertion Failure on Code Coverage and Spectrum-based Fault Localization
Md. Ashraf Uddin, Shaowei Wang, An Ran Chen +3
An assertion is commonly used to validate the expected programs behavior (e.g., if the returned value of a method equals an expected value) in software testing. Although it is a re…
Software Requirements Specification of the IUfA's UUIS -- a Team 2 COMP5541-W10 Project Approach
Omer Shahid Ahmad, Faisal Alrashdi, Jason +6
In the 52-page document, we describe our approach to the Software Requirements Specification of the IUfA's UUIS prototype. This includes the overall system description, functional…
Studying Duplicate Logging Statements and Their Relationships with Code Clones
Zhenhao Li, Tse-Hsun, Chen +2
In this paper, we focus on studying duplicate logging statements, which are logging statements that have the same static text message. We manually studied over 4K duplicate logging…
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
Harit Vishwakarma, Reid, Chen +4
Auto-labeling is an important family of techniques that produce labeled training sets with minimum manual labeling. A prominent variant, threshold-based auto-labeling (TBAL), works…
Crash Report Enhancement with Large Language Models: An Empirical Study
S M Farah Al Fahim, Md Nakhla Rafi, Zeyang Ma +3
Crash reports are central to software maintenance, yet many lack the diagnostic detail developers need to debug efficiently. We examine whether large language models can enhance cr…
Sparse random Hamiltonians are quantumly easy
Chi-Fang, Chen, Alexander M. Dalzell +3
A candidate application for quantum computers is to simulate the low-temperature properties of quantum systems. For this task, there is a well-studied quantum algorithm that perfor…
Studying and Benchmarking Large Language Models For Log Level Suggestion
Yi Wen Heng, Zeyang Ma, Zhenhao Li +3
Large Language Models (LLMs) have become a focal point of research across various domains, including software engineering, where their capabilities are increasingly leveraged. Rece…
Nucleon resonance production in the reaction
Jin-Quan Fan, Shao-Fei, Chen +1
In this work, we perform a study of nucleon resonance production in the reaction within an effective Lagrangian approach. In our model, we consider the excitation o…
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11
We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from th…
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Jiazhen Pan, Bailiang Jian, Paul Hager +19
The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…