papers

Publications (103)

cs.SE2026

From Horizontal Layering to Vertical Integration: A Comparative Study of the AI-Driven Software Development Paradigm

Chi Zhang, Zehan Li, Ziqian Zhong +4

This paper examines the organizational implications of Generative AI adoption in software engineering through a multiple-case comparative study. We contrast two development environ…

physics.acc-ph2026

Damping dynamics of the centroid oscillation of a relativistic laser pulse in a plasma channel

Yuhui Xia, Zhenan Wang, Ziyao Tang +10

The centroid oscillation of an offset laser pulse propagating in a preformed plasma channel is investigated through theoretical analysis and three-dimensional particle-in-cell simu…

cs.CV2020

Powering One-shot Topological NAS with Stabilized Share-parameter Proxy

Ronghao Guo, Chen Lin, Chuming Li +4

One-shot NAS method has attracted much interest from the research community due to its remarkable training efficiency and capacity to discover high performance models. However, the…

cs.IR2025

GENET: Unleashing the Power of Side Information for Recommendation via Hypergraph Pre-training

Yang Li, Qi'ao Zhao, Chen Lin +3

Recommendation with side information has drawn significant research interest due to its potential to mitigate user feedback sparsity. However, existing models struggle with general…

physics.optics2025

Observation of Arbitrarily Configurable Nonlinear Topological Modes

Kai Bai, Chen Lin, Jia-Zheng Li +3

Nonlinear topology is an emerging field that combines the intrinsic reconfigurability of nonlinear systems with the robustness of topological protection, offering fertile ground fo…

cs.CL2025

DeepThink: Aligning Language Models with Domain-Specific User Intents

Yang Li, Mingxuan Luo, Yeyun Gong +4

Supervised fine-tuning with synthesized instructions has been a common practice for adapting LLMs to domain-specific QA tasks. However, the synthesized instructions deviate from re…

physics.med-ph2021

Local vaccination and systemic tumor suppression via irradiation and manganese adjuvant in mice

Chunyang Lu, Jing Qian, Jianfeng Lv +15

Presently 4T-1 luc cells were irradiated with proton under ultra-high dose rate FLASH or with gamma-ray with conventional dose rate, and then subcutaneous vaccination with or witho…

cs.CV2021

PSViT: Better Vision Transformer via Token Pooling and Attention Sharing

Boyu Chen, Peixia Li, Baopu Li +6

In this paper, we observe two levels of redundancies when applying vision transformers (ViT) for image recognition. First, fixing the number of tokens through the whole network pro…

cs.CL2025

Rho-1: Not All Tokens Are What You Need

Zhenghao Lin, Zhibin Gou, Yeyun Gong +8

Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that "9l training". Our ini…

econ.GN2024

Measuring Gender and Racial Biases in Large Language Models

Jiafu An, Difang Huang, Chen Lin +1

In traditional decision making processes, social biases of human decision makers can lead to unequal economic outcomes for underrepresented social groups, such as women, racial or…

cs.IR2022

Visual Encoding and Debiasing for CTR Prediction

Si Chen, Chen Lin, Wanxian Guan +7

Extracting expressive visual features is crucial for accurate Click-Through-Rate (CTR) prediction in visual search advertising systems. Current commercial systems use off-the-shelf…

cond-mat.mtrl-sci2021

Two-dimensional Dirac nodal-line semimetal protected by symmetry

Xingxia Cui, Yafei Li, Deping Guo +12

Dirac nodal line semimetals (DNLSs) host relativistic quasiparticles in their one-dimensional (1D) Dirac nodal line (DNL) bands that are protected by certain crystalline symmetries…

cs.CV2022

UWC: Unit-wise Calibration Towards Rapid Network Compression

Chen Lin, Zheyang Li, Bo Peng +4

This paper introduces a post-training quantization~(PTQ) method achieving highly efficient Convolutional Neural Network~ (CNN) quantization with high performance. Previous PTQ meth…

math.NT2026

The -rationality of and

Chen Lin, Xuejun Guo

In this paper, we construct new families of imaginary and real quadratic fields that are -rational. In the imaginary case, we prove that for any positive integer and any int…

cs.DB2024

Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation

Yihang Zheng, Bo Li, Zhenghao Lin +6

The development of Large Language Models (LLMs) has revolutionized QA across various industries, including the database domain. However, there is still a lack of a comprehensive be…

physics.plasm-ph2015

Reduce proton energy spread by target ablation

Shuan Zhao, Chen Lin, Jiaer Chen +1

It's shown that, with strong target ablation monoenergetic protons along the laser direction is available during the laser aluminum foil interaction, which is different from the cl…

physics.med-ph2025

Dependence of the Radical Dynamics on the Beam Temporal Profile in FLASH Radiotherapy

Jianhan Sun, Xianghui Kong, Jianfeng Lv +6

Purpose: This study aims to investigate the impact of the beam temporal profile on the radical dynamics and inter-track interactions of FLASH radiotherapy, supporting parameter opt…

cs.IR2022

Shilling Black-box Recommender Systems by Learning to Generate Fake User Profiles

Chen Lin, Si Chen, Meifang Zeng +3

Due to the pivotal role of Recommender Systems (RS) in guiding customers towards the purchase, there is a natural motivation for unscrupulous parties to spoof RS for profits. In th…

cond-mat.soft2024

Active Liquid-Liquid Phase-Separation in a Confining Environment

Chen Lin, Robijn Bruinsma

Active liquid-liquid phase separation (LLPS) in a confining environment is believed to play an important role in cell biology. Recently, it was shown that when active noise at the…

physics.chem-ph2026

DFT Accuracy on Crystal Structure Prediction with Machine Learning Interatomic Potentials

Laurence I. Midgley, Chen Lin, J. Harry Moore +8

We present an evaluation of CSP-MACE-Ã , a machine learning interatomic potential intended to replace DFT in crystal structure prediction (CSP). We decompose the total energy into…

physics.app-ph2018

Positioning of Transparent Targets Using Defocusing Method in a Laser Proton Accelerator

Yinren Shou, Dahui Wang, Pengjie Wang +11

We report a positioning method for transparent targets with an accuracy of \SI{2}{μm} for a compact laser proton accelerator. The positioning system consists of two light-emitting…

physics.optics2026

Impact of noise on nonlinear-exceptional-point-based sensors

Kai Bai, Chen Lin, Meng Xiao

Nonlinear exceptional points (NEPs), a new type of spectral singularity in nonlinear non-Hermitian systems, are expected to address the noise divergence issue encountered at linear…

cs.LG2023

A Deep Reinforcement Learning Approach for Interactive Search with Sentence-level Feedback

Jianghong Zhou, Joyce C. Ho, Chen Lin +1

Interactive search can provide a better experience by incorporating interaction feedback from the users. This can significantly improve search accuracy as it helps avoid irrelevant…

cs.LG2024

Revealing Decurve Flows for Generalized Graph Propagation

Chen Lin, Liheng Ma, Yiyang Chen +3

This study addresses the limitations of the traditional analysis of message-passing, central to graph learning, by defining {\em \textbf{generalized propagation}} with directed and…

cs.CL2022

Unsupervised Extractive Summarization with Heterogeneous Graph Embeddings for Chinese Document

Chen Lin, Ye Liu, Siyu An +1

In the scenario of unsupervised extractive summarization, learning high-quality sentence representations is essential to select salient sentences from the input document. Previous…

cs.CV2022

Contrastive Graph Multimodal Model for Text Classification in Videos

Ye Liu, Changchong Lu, Chen Lin +2

The extraction of text information in videos serves as a critical step towards semantic understanding of videos. It usually involved in two steps: (1) text recognition and (2) text…

cs.CV2019

Improving One-shot NAS by Suppressing the Posterior Fading

Xiang Li, Chen Lin, Chuming Li +4

There is a growing interest in automated neural architecture search (NAS). To improve the efficiency of NAS, previous approaches adopt weight sharing method to force all models sha…

q-bio.QM2026

Empowering Chemical Structures with Biological Insights for Scalable Phenotypic Virtual Screening

Xiaoqing Lian, Pengsen Ma, Tengfeng Ma +9

Motivation: The scalable identification of bioactive compounds is essential for contemporary drug discovery. This process faces a key trade-off: structural screening offers scalabi…

stat.ME2026

Improve Power of Knockoffs with Annotation Information of Covariates

Xiangyu Zhang, Lijun Wang, Changjun Li +2

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these vari…

quant-ph2025

Efficient and Universal Neural-Network Decoder for Stabilizer-Based Quantum Error Correction

Gengyuan Hu, Wanli Ouyang, Chao-Yang Lu +2

Scaling quantum computing to practical applications necessitates reliable quantum error correction. Although numerous correction codes have been proposed, the overall correction ef…

cs.LG2018

Synaptic Strength For Convolutional Neural Network

Chen Lin, Zhao Zhong, Wei Wu +1

Convolutional Neural Networks(CNNs) are both computation and memory intensive which hindered their deployment in mobile devices. Inspired by the relevant concept in neural science…

quant-ph2024

On the distillablity conjecture in matrix theory

Saiqi Liu, Chen Lin

The distillability conjecture of two-copy 4 by 4 Werner states is one of the main open problems in quantum information. We prove two special cases of the conjecture. The first case…

cs.CV2021

DETR for Crowd Pedestrian Detection

Matthieu Lin, Chuming Li, Xingyuan Bu +5

Pedestrian detection in crowd scenes poses a challenging problem due to the heuristic defined mapping from anchors to pedestrians and the conflict between NMS and highly overlapped…

cs.LG2025

Large-scale automatic carbon ion treatment planning for head and neck cancers via parallel multi-agent reinforcement learning

Jueye Zhang, Chao Yang, Youfang Lai +9

Head-and-neck cancer (HNC) planning is difficult because multiple critical organs-at-risk (OARs) are close to complex targets. Intensity-modulated carbon-ion therapy (IMCT) offers…

physics.plasm-ph2025

First observation of shock waves induced by laser-accelerated proton beams

Yanlyu Fang, Xiaoyun Le, Yang Yan +5

We demonstrate, for the first time, that laser-accelerated protons can induce shock waves in materials. The ultra-short pulse width of laser-driven protons enables them to deposit…

cs.CV2019

AM-LFS: AutoML for Loss Function Search

Chuming Li, Yuan Xin, Chen Lin +4

Designing an effective loss function plays an important role in visual analysis. Most existing loss function designs rely on hand-crafted heuristics that require domain experts to…

cs.LG2020

Adaptive Gradient Method with Resilience and Momentum

Jie Liu, Chen Lin, Chuming Li +4

Several variants of stochastic gradient descent (SGD) have been proposed to improve the learning effectiveness and efficiency when training deep neural networks, among which some r…

cs.IR2026

Enhancing Healthcare Search Intent Recognition with Query Representation Learning and Session Context

Harshita Jagdish Sahijwani, Madhav Sigdel, Song Aslan +4

Classifying the intent behind healthcare search queries is crucial for improving the delivery of online healthcare information. The intricate nature of medical search queries, coup…

cond-mat.str-el2025

Revisiting the Broken Symmetry Phase of Solid Hydrogen: A Neural Network Variational Monte Carlo Study

Shengdu Chai, Chen Lin, Xinyang Dong +4

The crystal structure of high-pressure solid hydrogen remains a fundamental open problem. Although the research frontier has mostly shifted toward ultra-high pressure phases above…

cs.CV2021

Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture Search

Mingzhu Shen, Feng Liang, Ruihao Gong +6

Quantization Neural Networks (QNN) have attracted a lot of attention due to their high efficiency. To enhance the quantization accuracy, prior works mainly focus on designing advan…

cs.CV2021

Inception Convolution with Efficient Dilation Search

Jie Liu, Chuming Li, Feng Liang +5

As a variant of standard convolution, a dilated convolution can control effective receptive fields and handle large scale variance of objects without introducing additional computa…

cs.CL2026

QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs

Shupeng Li, Weipeng Lu, Linyun Liu +16

Domain-specific enhancement of Large Language Models (LLMs) within the financial context has long been a focal point of industrial application. While previous models such as Bloomb…

cs.CL2026

LFQA-E: Carefully Benchmarking Long-form QA Evaluation

Yuchen Fan, Chen Lin, Xin Zhong +11

Long-Form Question Answering (LFQA) involves generating comprehensive, paragraph-level responses to open-ended questions, which poses a significant challenge for evaluation due to…

cs.CV2026

Efficient Token Pruning for LLaDA-V

Zhewen Wan, Tianchen Song, Chen Lin +2

Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional at…

cs.LG2026

Delayed Feedback Modeling for Post-Click Gross Merchandise Volume Prediction: Benchmark, Insights and Approaches

Xinyu Li, Sishuo Chen, Guipeng Xv +7

The prediction objectives of online advertisement ranking models are evolving from probabilistic metrics like conversion rate (CVR) to numerical business metrics like post-click gr…

math.NT2024

The asymptotic estimation of prime ideals in imaginary quadratic fields and Chebyshev's bias

Chen Lin, Chenhao Tang, Xuejun Guo

We study the asymptotic estimation of prime ideals that satisfy certain congruence and argument conditions in imaginary quadratic fields. We also discuss the phenomenon of Chebyshe…

math.NT2026

On the Fractional Parts of Polynomials Modulo

Xuejun Guo, Chen Lin, Zhefeng Xu

We study a half-interval distribution problem for polynomial residues modulo an odd prime : how often the fractional part of lies in the upper half of the unit interva…

cs.MA2025

From Intention To Implementation: Automating Biomedical Research via LLMs

Yi Luo, Linghang Shi, Yihao Li +4

Conventional biomedical research is increasingly labor-intensive due to the exponential growth of scientific literature and datasets. Artificial intelligence (AI), particularly Lar…

cs.IR2024

DocReLM: Mastering Document Retrieval with Language Model

Gengchen Wei, Xinle Pang, Tianning Zhang +5

With over 200 million published academic documents and millions of new documents being written each year, academic researchers face the challenge of searching for information withi…

cs.CV2019

Online Hyper-parameter Learning for Auto-Augmentation Strategy

Chen Lin, Minghao Guo, Chuming Li +5

Data augmentation is critical to the success of modern deep learning techniques. In this paper, we propose Online Hyper-parameter Learning for Auto-Augmentation (OHL-Auto-Aug), an…

cs.CL2025

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models

Zhenghao Lin, Zihao Tang, Xiao Liu +31

We introduce Sigma, an efficient large language model specialized for the system domain, empowered by a novel architecture including DiffQKV attention, and pre-trained on our metic…

cs.CV2020

PV-NAS: Practical Neural Architecture Search for Video Recognition

Zihao Wang, Chen Lin, Lu Sheng +2

Recently, deep learning has been utilized to solve video recognition problem due to its prominent representation ability. Deep neural networks for video tasks is highly customized…

cs.CV2022

Fast-MoCo: Boost Momentum-based Contrastive Learning with Combinatorial Patches

Yuanzheng Ci, Chen Lin, Lei Bai +1

Contrastive-based self-supervised learning methods achieved great success in recent years. However, self-supervision requires extremely long training epochs (e.g., 800 epochs for M…

cs.CL2025

HiCaM: A Hierarchical-Causal Modification Framework for Long-Form Text Modification

Yuntao Shi, Yi Luo, Yeyun Gong +1

Large Language Models (LLMs) have achieved remarkable success in various domains. However, when handling long-form text modification tasks, they still face two major problems: (1)…

cs.CV2021

Evolving Search Space for Neural Architecture Search

Yuanzheng Ci, Chen Lin, Ming Sun +3

The automation of neural architecture design has been a coveted alternative to human experts. Recent works have small search space, which is easier to optimize but has a limited up…

cs.CV2019

Computation Reallocation for Object Detection

Feng Liang, Chen Lin, Ronghao Guo +4

The allocation of computation resources in the backbone is a crucial issue in object detection. However, classification allocation pattern is usually adopted directly to object det…

cs.CL2023

Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph Denoise

Zhenghao Lin, Yeyun Gong, Yelong Shen +5

In this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pretrained diffusion language m…

cs.SE2021

Improving Code Summarization with Block-wise Abstract Syntax Tree Splitting

Chen Lin, Zhichao Ouyang, Junqing Zhuang +3

Automatic code summarization frees software developers from the heavy burden of manual commenting and benefits software development and maintenance. Abstract Syntax Tree (AST), whi…

physics.chem-ph2025

Cross Learning between Electronic Structure Theories for Unifying Molecular, Surface, and Inorganic Crystal Foundation Force Fields

Ilyes Batatia, Chen Lin, Joseph Hart +5

Creating a single unified interatomic potential capable of attaining ab initio accuracy across all chemistry remains a long-standing challenge in computational chemistry and materi…

cs.CL2025

Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs

Yi Luo, Qiwen Wang, Junqi Yang +5

Generalized Category Discovery (GCD) aims to classify both known and novel categories using partially labeled data that contains only known classes. Despite achieving strong perfor…

cs.IR2023

PROD: Progressive Distillation for Dense Retrieval

Zhenghao Lin, Yeyun Gong, Xiao Liu +8

Knowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the s…

quant-ph2023

Low-overhead pieceable fault-tolerant construction of logical controlled-phase circuit for degenerate quantum code

Chen Lin, Guowu Yang

We designed an search algorithm in order to find a non-transversal but fault-tolerant construction of a logical controlled-phase gate for general [[n,1,d]] degenerate quantum code.…

cs.CV2022

Efficient Joint-Dimensional Search with Solution Space Regularization for Real-Time Semantic Segmentation

Peng Ye, Baopu Li, Tao Chen +6

Semantic segmentation is a popular research topic in computer vision, and many efforts have been made on it with impressive results. In this paper, we intend to search an optimal n…

nucl-ex2024

Ultra-short lifetime isomer studies from photonuclear reactions using laser-driven ultra-intense γ-ray

Di Wu, Haoyang Lan, Jiaxing Liu +29

Isomers, ubiquitous populations of relatively long-lived nuclear excited states, play a crucial role in nuclear physics. However, isomers with half-life times of several seconds or…

cs.LG2020

Improving Auto-Augment via Augmentation-Wise Weight Sharing

Keyu Tian, Chen Lin, Ming Sun +3

The recent progress on automatically searching augmentation policies has boosted the performance substantially for various tasks. A key component of automatic augmentation search i…

cs.LG2026

Modeling Cascaded Delay Feedback for Online Net Conversion Rate Prediction: Benchmark, Insights and Solutions

Mingxuan Luo, Guipeng Xv, Sishuo Chen +8

In industrial recommender systems, conversion rate (CVR) is widely used for traffic allocation, but it fails to fully reflect recommendation effectiveness because it ignores refund…

eess.IV2022

SuperVessel: Segmenting High-resolution Vessel from Low-resolution Retinal Image

Yan Hu, Zhongxi Qiu, Dan Zeng +3

Vascular segmentation extracts blood vessels from images and serves as the basis for diagnosing various diseases, like ophthalmic diseases. Ophthalmologists often require high-reso…

cs.IR2020

FLEN: Leveraging Field for Scalable CTR Prediction

Wenqiang Chen, Lizhang Zhan, Yuanlong Ci +3

Click-Through Rate (CTR) prediction has been an indispensable component for many industrial applications, such as recommendation systems and online advertising. CTR prediction syst…

physics.chem-ph2025

A foundation model for atomistic materials chemistry

Ilyes Batatia, Philipp Benner, Yuan Chiang +85

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much…

physics.app-ph2021

Efficient light-emitting diodes based on oriented perovskite nanoplatelets

Jieyuan Cui, Yang Liu, Yunzhou Deng +15

Solution-processed planar perovskite light-emitting diodes (LEDs) promise high-performance and cost-effective electroluminescent (EL) devices ideal for large-area display and light…

cs.CL2024

Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph

Jiashuo Sun, Chengjin Xu, Lumingyuan Tang +6

Although large language models (LLMs) have achieved significant success in various tasks, they often struggle with hallucination problems, especially in scenarios requiring deep an…

stat.CO2024

HDTSA: An R package for high-dimensional time series analysis

Jinyuan Chang, Jing He, Chen Lin +1

High-dimensional time series analysis has become increasingly important in fields such as finance, economics, and biology. The two primary tasks for high-dimensional time series an…

cs.CV2023

Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling

Keyu Tian, Yi Jiang, Qishuai Diao +3

We identify and overcome two key obstacles in extending the success of BERT-style pre-training, or the masked image modeling, to convolutional networks (convnets): (i) convolution…

cs.CL2024

AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators

Xingwei He, Zhenghao Lin, Yeyun Gong +7

Many natural language processing (NLP) tasks rely on labeled data to train machine learning models with high performance. However, data annotation is time-consuming and expensive,…

cs.CL2024

Competition-Level Problems are Effective LLM Evaluators

Yiming Huang, Zhenghao Lin, Xiao Liu +8

Large language models (LLMs) have demonstrated impressive reasoning capabilities, yet there is ongoing debate about these abilities and the potential data contamination problem rec…

math.NT2026

Counting Polynomial-type Exceptional Units on Algebraic Varieties over Number Fields

Chen Lin, Kaihan Tang

Previous research on exceptional units has primarily focused on the ring of rational integers or abstract finite rings, often restricted to linear or quadratic constraints. In this…

physics.plasm-ph2026

Impact of Residual Angular Chirp in a Petawatt-class Laser System on Laser-driven Proton Acceleration

Qingfan Wu, Minjian Wu, Jiarui Zhao +25

The paper shows that a small residual angular chirp caused by misaligned grating compressors in a petawatt laser degrades the focal spot and reduces proton energies, and that remov…

#laser-driven proton acceleration#angular chirp#petawatt laser#focal spot quality
cs.CV2021

GLiT: Neural Architecture Search for Global and Local Image Transformer

Boyu Chen, Peixia Li, Chuming Li +6

We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones a…

cs.CL2024

Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models

Yi Luo, Zhenghao Lin, Yuhao Zhang +7

Large Language Models (LLMs) exhibit impressive capabilities but also present risks such as biased content generation and privacy issues. One of the current alignment techniques in…

cs.CV2024

LOCR: Location-Guided Transformer for Optical Character Recognition

Yu Sun, Dongzhan Zhou, Chen Lin +3

Academic documents are packed with texts, equations, tables, and figures, requiring comprehensive understanding for accurate Optical Character Recognition (OCR). While end-to-end O…

cs.CL2026

From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation

Jiahao Wang, Weiyu Xie, Mingxing Zhang +10

Retrieval-Augmented Generation enhances Large Language Models by integrating external knowledge, which reduces hallucinations but increases prompt length. This increase leads to hi…

cs.CL2022

Sentiment-Aware Word and Sentence Level Pre-training for Sentiment Analysis

Shuai Fan, Chen Lin, Haonan Li +6

Most existing pre-trained language representation models (PLMs) are sub-optimal in sentiment analysis tasks, as they capture the sentiment information from word-level while under-c…

cs.CV2021

BN-NAS: Neural Architecture Search with Batch Normalization

Boyu Chen, Peixia Li, Baopu Li +5

We present BN-NAS, neural architecture search with Batch Normalization (BN-NAS), to accelerate neural architecture search (NAS). BN-NAS can significantly reduce the time required b…

cs.CV2023

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Ziyi Lin, Chris Liu, Renrui Zhang +13

We present SPHINX, a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, tuning tasks, and visual embeddings. First, for stronger vision-languag…

cs.CL2024

APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning

Jiashuo Sun, Hang Zhang, Chen Lin +3

Long-form numerical reasoning in financial analysis aims to generate a reasoning program to calculate the correct answer for a given question. Previous work followed a retriever-ge…

cs.AI2021

Sequential Recommendation in Online Games with Multiple Sequences, Tasks and User Levels

Si Chen, Yuqiu Qian, Hui Li +1

Online gaming is growing faster than ever before, with increasing challenges of providing better user experience. Recommender systems (RS) for online games face unique challenges s…

physics.optics2026

HotLoop Optimization of Petawatt Laser Focal Spot via a Twin-Focus Scheme

Qingfan Wu, Ying Gao, Minjian Wu +25

Achieving diffraction-limited focusing of high-power laser pulses to generate ultra-high intensities is crucial for developing compact laser-driven particle accelerators and explor…

cs.LG2025

Implicit Neural Representations for Chemical Reaction Paths

Kalyan Ramakrishnan, Lars L. Schaaf, Chen Lin +2

We show that neural networks can be optimized to represent minimum energy paths as continuous functions, offering a flexible alternative to discrete path-search methods such as Nud…

cs.IR2025

Improving Multi-modal Recommender Systems by Denoising and Aligning Multi-modal Content and User Feedback

Guipeng Xv, Xinyu Li, Ruobing Xie +5

Multi-modal recommender systems (MRSs) are pivotal in diverse online web platforms and have garnered considerable attention in recent years. However, previous studies overlook the…

math.NT2026

Matrix Kloosterman sums and product-trace estimates for semisimple algebras

Xuejun Guo, Chen Lin, Chenhao Tang

Let , and . For , and , let be the number of $r…

cs.CL2024

Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models

Jiashuo Sun, Yi Luo, Yeyun Gong +4

Large language models (LLMs) can achieve highly effective performance on various reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting as demonstrations. H…

stat.ME2017

Group-Average and Convex Clustering for Partially Heterogeneous Linear Regression

Lu Lin, Jun Lu, Chen Lin

In this paper, a subgroup least squares and a convex clustering are introduced for inferring a partially heterogenous linear regression that has potential application in the areas…

cs.RO2025

An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework

Zeyu Wang, Frank P. -W. Lo, Qian Chen +7

Existing LLM-enabled multi-agent frameworks are predominantly limited to digital or simulated environments and confined to narrowly focused knowledge domain, constraining their app…

cs.LG2024

Self-consistent Validation for Machine Learning Electronic Structure

Gengyuan Hu, Gengchen Wei, Zekun Lou +4

Machine learning has emerged as a significant approach to efficiently tackle electronic structure problems. Despite its potential, there is less guarantee for the model to generali…

cs.IR2026

ThinkRec: Thinking-based recommendation via LLM

Qihang Yu, Kairui Fu, Zheqi Lv +6

Recent advances in large language models (LLMs) have enabled more semantic-aware recommendations through natural language generation. Existing LLM for recommendation (LLM4Rec) meth…

cs.LG2026

Scalable Equilibrium Sampling with Sequential Boltzmann Generators

Charlie B. Tan, Avishek Joey Bose, Chen Lin +3

Scalable sampling of molecular states in thermodynamic equilibrium is a long-standing challenge in statistical physics. Boltzmann generators tackle this problem by pairing normaliz…

cs.CV2024

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Peng Gao, Le Zhuo, Dongyang Liu +17

Sora unveils the potential of scaling Diffusion Transformer for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lac…

cs.LG2022

Detecting Elevated Air Pollution Levels by Monitoring Web Search Queries: Deep Learning-Based Time Series Forecasting

Chen Lin, Safoora Yousefi, Elvis Kahoro +4

Real-time air pollution monitoring is a valuable tool for public health and environmental surveillance. In recent years, there has been a dramatic increase in air pollution forecas…

cs.LG2026

FAAR: Format-Aware Adaptive Rounding for NVFP4

Hanglin Li, Shuchang Tian, Chen Lin +2

Deploying large language models (LLMs) on edge devices requires extremely low-bit quantization. Ultra-low precision formats such as NVFP4 offer a promising solution for reducing me…

cs.LG2026

ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation

Chen Lin, Kedi Chen, Wei Zhang

On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but standard OPD treats all student-generated outputs (SGOs) equally r…