papers

Publications (118)

gr-qc2021

A Dark Energy model from Generalized Proca Theory

Chao-Qiang Geng, Yan-Ting Hsu, Jhih-Rong Lu +1

We consider a specific dark energy model, which only includes the Lagrangian up to the cubic order in terms of the vector field self-interactions in the generalized Proca theory. W…

cs.CV2026

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

Hanqi Yan, Xiangxiang Cui, Lu Yin +4

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing align…

cs.LG2025

Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing

Zihao Wang, Enneng Yang, Lu Yin +2

Model merging leverages multiple finetuned expert models to construct a multi-task model with low cost, and is gaining increasing attention. However, as a growing number of finetun…

cs.CL2025

TODO: Enhancing LLM Alignment with Ternary Preferences

Yuxiang Guo, Lu Yin, Bo Jiang +1

Aligning large language models (LLMs) with human intent is critical for enhancing their performance across a variety of tasks. Standard alignment techniques, such as Direct Prefere…

cs.CV2025

MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic Disturbances

Yunzhe Shao, Xinyu Yi, Lu Yin +3

This paper proposes a novel method called MagShield, designed to address the issue of magnetic interference in sparse inertial motion capture (MoCap) systems. Existing Inertial Mea…

astro-ph.CO2021

Does Hubble Tension Signal a Breakdown in FLRW Cosmology?

Chethan Krishnan, Roya Mohayaee, Eoin Ó Colgáin +2

The tension between early and late Universe probes of the Hubble constant has motivated various new FLRW cosmologies. Here, we reanalyse the Hubble tension with a recent age of the…

astro-ph.CO2020

Modified Cosmology Models from Thermodynamical Approach

Chao-Qiang Geng, Yan-Ting Hsu, Jhih-Rong Lu +1

We apply the first law of thermodynamics to the apparent horizon of the universe with the power-law corrected and non-extensive Tsallis entropies rather than the Bekenstein-Hawking…

cs.LG2024

Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients

Zhenyu Zhang, Ajay Jaiswal, Lu Yin +4

Training Large Language Models (LLMs) is memory-intensive due to the large number of parameters and associated optimization states. GaLore, a recent method, reduces memory usage by…

astro-ph.CO2022

On Larger Values in the CMB Dipole Direction

Orlando Luongo, Marco Muccino, Eoin Ó Colgáin +2

On the assumption that quasars (QSO) and gamma-ray bursts (GRB) represent \textit{standardisable candles}, we provide evidence that the Hubble constant adopts larger values i…

astro-ph.CO2021

Observational Constraints on the Cosmology with Holographic Dark Fluid

Da Huang, Bum-Hoon Lee, Gansukh Tumurtushaa +2

We consider the holographic Friedman-Robertson-Walker (hFRW) universe on the 4-dimensional membrane embedded in the 5-dimensional bulk spacetime and fit the parameters with the obs…

physics.optics2025

SMILE: a universal tool for modulated-enhanced localization microscopy to achieve minimal three-dimensional resolution

Hongfei Zhu, Yile Sun, Xinxun Yang +8

Modulation-enhanced localization microscopy (MELM) has demonstrated significant improvements in both lateral and axial localization precision compared to conventional single-molecu…

cs.CV2022

Semantic-Based Few-Shot Learning by Interactive Psychometric Testing

Lu Yin, Vlado Menkovski, Yulong Pei +1

Few-shot classification tasks aim to classify images in query sets based on only a few labeled examples in support sets. Most studies usually assume that each image in a task has a…

cs.CV2020

DymSLAM:4D Dynamic Scene Reconstruction Based on Geometrical Motion Segmentation

Chenjie Wang, Bin Luo, Yun Zhang +6

Most SLAM algorithms are based on the assumption that the scene is static. However, in practice, most scenes are dynamic which usually contains moving objects, these methods are no…

astro-ph.CO2020

Constraints on a special running vacuum model

Chao-Qiang Geng, Chung-Chi Lee, Lu Yin

We study a special running vacuum model (RVM) with , where , and are the model parameters and is the Hubble one. This RVM has…

cs.AI2026

Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution

Qiao Xiao, Alan Ansell, Boqian Wu +4

Sparse large language models (LLMs) offer an attractive direction toward efficient deployment, but adapting them to downstream tasks remains challenging. The central difficulty is…

astro-ph.CO2023

Is Cosmic Birefringence model-dependent?

Lu Yin, Joby Kochappan, Tuhin Ghosh +1

Exciting clues to isotropic cosmic birefringence have recently been detected in the cross-power spectra of the polarization data of the cosmic microwave background (CMB). Earl…

cs.LG2021

Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training

Shiwei Liu, Lu Yin, Decebal Constantin Mocanu +1

In this paper, we introduce a new perspective on training deep neural networks capable of state-of-the-art performance without the need for the expensive over-parameterization by p…

cs.CL2026

Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving

Xin Xu, Yan Xu, Tianhao Chen +9

Existing approaches to mathematical reasoning with large language models (LLMs) rely on Chain-of-Thought (CoT) for generalizability or Tool-Integrated Reasoning (TIR) for precise c…

cs.CL2024

CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization

Zheyan Qu, Lu Yin, Zitong Yu +2

Large language models (LLMs) have demonstrated astonishing capabilities in natural language processing (NLP) tasks, sparking interest in their application to professional domains w…

cs.LG2026

The Curse of Depth in Large Language Models

Wenfang Sun, Xinyuan Song, Pengxiang Li +3

In this paper, we introduce the Curse of Depth, a concept that highlights, explains, and addresses the recent observation in modern Large Language Models (LLMs) where nearly half o…

eess.SP2019

Design and Performance Analysis of Multi-scale NOMA for 5G Positioning

Lu Yin, Jiameng Cao, Zhongliang Deng +4

This paper presents a feasibility study for a novel positioning-communication integrated signal called Multi-Scale Non-Orthogonal Multiple Access (MS-NOMA) for 5G positioning. One…

cs.LG2026

Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs

Lu Yin, Ajay Jaiswal, Shiwei Liu +2

We present Junk DNA Hypothesis by adopting a novel task-centric angle for the pre-trained weights of large language models (LLMs). It has been believed that weights in LLMs contain…

cs.CL2024

Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning

Abhinav Bandari, Lu Yin, Cheng-Yu Hsieh +5

Network pruning has emerged as a potential solution to make LLMs cheaper to deploy. However, existing LLM pruning approaches universally rely on the C4 dataset as the calibration d…

cs.LG2026

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

Di He, Songjun Tu, Keyu Wang +2

Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural…

hep-ph2023

WIMPs in Dilatonic Einstein Gauss-Bonnet Cosmology

Anirban Biswas, Arpan Kar, Bum-Hoon Lee +5

We use the Weakly Interacting Massive Particle (WIMP) thermal decoupling scenario to probe Cosmologies in dilatonic Einstein Gauss-Bonnet (dEGB) gravity, where the Gauss-Bonnet ter…

cs.CL2025

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

Xin Xu, Tianhao Chen, Fan Zhang +11

While slow-thinking large language models (LLMs) exhibit reflection-like reasoning, commonly referred to as the "aha moment:, their ability to generate informative critiques and re…

cs.CV2025

E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation

Boqian Wu, Qiao Xiao, Shiwei Liu +5

Deep neural networks have evolved as the leading approach in 3D medical image segmentation due to their outstanding performance. However, the ever-increasing model size and computa…

physics.app-ph2020

High Performance Printed AgO-Zn Rechargeable Battery for Flexible Electronics

Lu Yin, Jonathan Scharf, Jessica Ma +8

The rise of flexible electronics calls for cost-effective and scalable batteries with good mechanical and electrochemical performance. In this work, we developed printable, polymer…

cs.LG2024

You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets

Tianjin Huang, Tianlong Chen, Meng Fang +8

Recent works have impressively demonstrated that there exists a subnetwork in randomly initialized convolutional neural networks (CNNs) that can match the performance of the fully…

cs.LG2024

Robust Active Learning (RoAL): Countering Dynamic Adversaries in Active Learning with Elastic Weight Consolidation

Ricky Maulana Fajri, Yulong Pei, Lu Yin +1

Despite significant advancements in active learning and adversarial attacks, the intersection of these two fields remains underexplored, particularly in developing robust active le…

cs.LG2026

Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization

Xinchen Han, Hossam Afifi, Michel Marot +2

Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency without proportional performance g…

cs.LG2026

LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning

Zihang Liu, Tianyu Pang, Oleg Balabanov +5

Recent studies have shown that supervised fine-tuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabilities. However, full fine-tuning (Full FT…

astro-ph.CO2023

Inevitable manifestation of wiggles in the expansion of the late Universe

Ozgur Akarsu, Eoin O. Colgain, Emre Ozulker +2

Using the fact that the comoving angular diameter distance to the last scattering surface is strictly constrained almost model independently, we show that, for any model agreeing w…

astro-ph.CO2022

Hints of FLRW Breakdown from Supernovae

Chethan Krishnan, Roya Mohayaee, Eoin Ó Colgáin +2

A 10\% difference in the scale for the Hubble parameter constitutes a clear problem for cosmology. Here, considering angular distribution of Type Ia supernovae (SN) within the Pant…

cs.CV2026

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

Shaojie Zhuang, Lu Yin, Guangshun Wei +3

Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing approaches rely on task-specific…

cs.LG2025

From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Ajay Jaiswal, Yifan Wang, Lu Yin +6

Large Language Models' (LLMs) weight matrices can often be expressed in low-rank form with potential to relax memory and compute resource requirements. Unlike prior efforts that fo…

cs.CV2023

Are Large Kernels Better Teachers than Transformers for ConvNets?

Tianjin Huang, Lu Yin, Zhenyu Zhang +5

This paper reveals a new appeal of the recently emerged large-kernel Convolutional Neural Networks (ConvNets): as the teacher in Knowledge Distillation (KD) for small-kernel ConvNe…

cs.AI2025

Into the Unknown: Applying Inductive Spatial-Semantic Location Embeddings for Predicting Individuals' Mobility Beyond Visited Places

Xinglei Wang, Tao Cheng, Stephen Law +6

Predicting individuals' next locations is a core task in human mobility modelling, with wide-ranging implications for urban planning, transportation, public policy and personalised…

cs.CV2025

Pushing the Limits of Sparsity: A Bag of Tricks for Extreme Pruning

Andy Li, Aiden Durrant, Milan Markovic +5

Pruning of deep neural networks has been an effective technique for reducing model size while preserving most of the performance of dense networks, crucial for deploying models on…

cs.GR2025

Improving Sparse IMU-based Motion Capture with Motion Label Smoothing

Zhaorui Meng, Lu Yin, Yangqing Hou +3

Sparse Inertial Measurement Units (IMUs) based human motion capture has gained significant momentum, driven by the adaptation of fundamental AI tools such as recurrent neural netwo…

cs.CL2024

Dynamic Data Pruning for Automatic Speech Recognition

Qiao Xiao, Pingchuan Ma, Adriana Fernandez-Lopez +7

The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitivel…

cs.AI2025

Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework

Zhuo Zhi, Chen Feng, Adam Daneshmend +6

Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic fr…

astro-ph.CO2022

Reducing the Tension with Exponential Acoustic Dark Energy

Lu Yin

The Hubble tension arises from different observations between the late-time and early Universe. We explore a new model with dark fluid, called the exponential Acoustic Dark Energy…

astro-ph.CO2025

Constraints on Cosmic Birefringence from SPIDER, Planck, and ACT observations

Lu Yin, Shuhang Xiong, Joby Kochappan +2

The Early Dark Energy (EDE) model has been proposed as a candidate mechanism to generate cosmic birefringence through a Chern-Simons coupling between a dynamical scalar field and t…

astro-ph.CO2024

Observational evidence for Early Dark Energy as a unified explanation for Cosmic Birefringence and the Hubble tension

Joby Kochappan, Lu Yin, Bum-Hoon Lee +1

We test the =3 Ultralight Axion-like model of Early Dark Energy (EDE) with the observationsof the mode of the cosmic microwave background (CMB) radiation, and local expansi…

astro-ph.CO2025

Does Gauss-Bonnet Inflationary Gravitational Waves satisfy the Pulsar Timing Arrays observations?

Lu Yin

The observations from pulsar timing arrays (PTAs), led by the North American Nanohertz Observatory for Gravitational Waves (NANOGrav), have provided opportunities to constrain prim…

cs.IR2021

Linear-Time Self Attention with Codeword Histogram for Efficient Recommendation

Yongji Wu, Defu Lian, Neil Zhenqiang Gong +4

Self-attention has become increasingly popular in a variety of sequence modeling tasks from natural language processing to recommendation, due to its effectiveness. However, self-a…

cs.LG2022

Sparse Training via Boosting Pruning Plasticity with Neuroregeneration

Shiwei Liu, Tianlong Chen, Xiaohan Chen +7

Works on lottery ticket hypothesis (LTH) and single-shot network pruning (SNIP) have raised a lot of attention currently on post-training pruning (iterative magnitude pruning), and…

hep-ph2025

Light quark energy loss in the flavor-dependent systems from holography

Le Zhang, Lu Yin, Guo-Dong Zhou +2

Using the holographic model of finite-endpoint-momentum shooting string approach, we study the instantaneous energy loss of light quarks for the flavor-dependent systems with $N_f…

cs.CV2026

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

Xiangxiang Cui, Tianjin Huang, Yifang Wang +2

Medical foundation models have achieved remarkable clinical performance, yet their robustness under real-world perturbations remains underexplored. We present a robustness benchmar…

cs.LG2025

Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models

Zihan Wang, Rui Pan, Jiarui Yao +7

We propose Chain-of-Experts (CoE), a new Mixture-of-Experts (MoE) architecture that introduces sequential expert communication within each layer. Unlike traditional MoE models, whe…

hep-th2024

Late-time Cosmology without Dark Sector but with Closed String Massless Sector

Hocheol Lee, Jeong-Hyuck Park, Liliana Velasco-Sevilla +1

We explore the possibility of solving the dark energy and the coincidence problems by postulating the massless sector of closed strings. This sector constitutes the gravitational m…

cs.LG2025

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Pengxiang Li, Lu Yin, Shiwei Liu

Large Language Models (LLMs) have achieved remarkable success, yet recent findings reveal that their deeper layers often contribute minimally and can be pruned without affecting ov…

cs.CL2026

SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment

Ziyang Chen, Zhenxuan Huang, Yile Wang +3

Traditional sentence embedding methods employ token-level contrastive learning on non-generative pre-trained models. Recently, there have emerged embedding methods based on generat…

astro-ph.CO2025

Do high redshift QSOs and GRBs corroborate JWST?

Eoin Ó Colgáin, M. M. Sheikh-Jabbari, Lu Yin

The James Webb Space Telescope (JWST) is reporting massive high redshift galaxies that appear challenging from the CDM perspective. Interpreted as a cosmological problem, this…

cs.LG2023

A Structural-Clustering Based Active Learning for Graph Neural Networks

Ricky Maulana Fajri, Yulong Pei, Lu Yin +1

In active learning for graph-structured data, Graph Neural Networks (GNNs) have shown effectiveness. However, a common challenge in these applications is the underutilization of cr…

cs.CV2021

Hierarchical Semantic Segmentation using Psychometric Learning

Lu Yin, Vlado Menkovski, Shiwei Liu +1

Assigning meaning to parts of image data is the goal of semantic image segmentation. Machine learning methods, specifically supervised learning is commonly used in a variety of tas…

cs.CV2026

CPR: Chained Perceptual Refinement for Coarse-to-Fine Medical Image Classification

Si-Yuan Lu, Hanruo Zhu, Ziquan Zhu +6

High resolution medical images contain fine grained, spatially sparse cues that are critical for diagnosis, yet preserving full resolution incurs substantial computational and memo…

astro-ph.CO2026

How much has DESI dark energy evolved since DR1?

Eoin Ó Colgáin, Saeed Pourojaghi, M. M. Sheikh-Jabbari +1

DESI has reported a dynamical dark energy (DE) signal based on the CDM model that is in conflict with Hubble tension. Recalling that the combination of DESI DR1 BAO and DR…

cs.CL2024

FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping

Ajay Jaiswal, Bodun Hu, Lu Yin +4

Autoregressive Large Language Models (e.g., LLaMa, GPTs) are omnipresent achieving remarkable success in language understanding and generation. However, such impressive capability…

gr-qc2022

Gravitational waves from the vacuum decay with LISA

Bum-Hoon Lee, Wonwoo Lee, Dong-han Yeom +1

We investigate the gravitational wave spectrum resulted from the cosmological first-order phase transition. We compare two models; one is a scalar field model without gravitation,…

cs.IR2021

Rethinking Lifelong Sequential Recommendation with Incremental Multi-Interest Attention

Yongji Wu, Lu Yin, Defu Lian +4

Sequential recommendation plays an increasingly important role in many e-commerce services such as display advertisement and online shopping. With the rapid development of these se…

astro-ph.CO2025

The CosmoVerse White Paper: Addressing observational tensions in cosmology with systematics and fundamental physics

Eleonora Di Valentino, Jackson Levi Said, Adam Riess +533

The standard model of cosmology has provided a good phenomenological description of a wide range of observations both at astrophysical and cosmological scales for several decades.…

cond-mat.mtrl-sci2021

Investigating Degradation Modes in Zn-AgO Aqueous Batteries with X-ray Micro Computed Tomography

Jonathan Scharf, Lu Yin, Christopher Redquest +7

To meet growing energy demands, degradation mechanisms of energy storage devices must be better understood. As a non-destructive tool, X-ray Computed Tomography (CT) has been incre…

cs.LG2025

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs

Keyu Wang, Tian Lyu, Guinan Su +4

Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing methods demonstrate strong performance reten…

cs.AI2026

Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning

Xinyuan Song, Keyu Wang, PengXiang Li +2

Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance…

cs.LG2020

Knowledge Elicitation using Deep Metric Learning and Psychometric Testing

Lu Yin, Vlado Menkovski, Mykola Pechenizkiy

Knowledge present in a domain is well expressed as relationships between corresponding concepts. For example, in zoology, animal species form complex hierarchies; in genomics, the…

astro-ph.CO2021

Can dark energy be dynamical?

Eoin Ó Colgáin, M. M. Sheikh-Jabbari, Lu Yin

We highlight shortcomings of the dynamical dark energy (DDE) paradigm. For parametric models with equation of state (EOS), for a given function of redshift…

cs.CV2024

Aspect-Based Few-Shot Learning

Tim van Engeland, Lu Yin, Vlado Menkovski

We generalize the formulation of few-shot learning by introducing the concept of an aspect. In the traditional formulation of few-shot learning, there is an underlying assumption t…

astro-ph.CO2026

Joint constraints on cosmic birefringence and early dark energy from ACT, Planck, DESI, and PantheonPlus

Lu Yin, Guo-Hong Du, Tian-Nuo Li +1

With the increasing number of high-precision astronomical observations, physical quantities that were previously inaccessible to accurate calculations, such as cosmic birefringence…

astro-ph.CO2018

Cosmological constraints on CDM models with time-varying fine structure constant

Jin-Jun Zhang, Lu Yin, Chao-Qiang Geng

We study the CDM models with being a function of the time-varying fine structure constant . We give a close look at the constraints on two specific CDM…

cs.AI2024

Multimodal Contrastive Learning of Urban Space Representations from POI Data

Xinglei Wang, Tao Cheng, Stephen Law +3

Existing methods for learning urban space representations from Point-of-Interest (POI) data face several limitations, including issues with geographical delineation, inadequate spa…

hep-ph2019

Stochastic Gravitational Waves from Inflaton Decays

Da Huang, Lu Yin

Due to the universality of gravitational interactions, it is generally expected that a stochastic gravitational wave (GW) background could form during the reheating period when the…

cs.LG2026

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity

Jiaxi Li, Lu Yin, Li Shen +5

Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck for widespread adoption. Lo…

cs.LG2025

Sebra: Debiasing Through Self-Guided Bias Ranking

Adarsh Kappiyath, Abhra Chaudhuri, Ajay Jaiswal +4

Ranking samples by fine-grained estimates of spuriosity (the degree to which spurious cues are present) has recently been shown to significantly benefit bias mitigation, over the t…

cs.CL2026

Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Models

Mingyu Cao, Alvaro H. C. Correia, Christos Louizos +2

Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, repeatedly deciding which positions to commit at each step. Standard decoding follows a g…

cs.LG2025

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

Tianhao Chen, Xin Xu, Zijing Liu +12

Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretra…

cs.SD2024

Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models

Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin +2

This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viabi…

astro-ph.CO2020

Running vacuum model in non-flat universe

Chao-Qiang Geng, Yan-Ting Hsu, Lu Yin +1

We investigate observational constraints on the running vacuum model (RVM) of in the spatially curved universe, where is the model parameter, cor…

cs.AI2026

AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

Zahratu Shabrina, Muhammad Asa, Jin Rui +2

The paper evaluates how state‑of‑the‑art vision‑language models (e.g., GPT‑4o, Claude 3.5 Sonnet, Gemini 2.0 Flash) predict building typology attributes from Google Street View ima…

#building typology classification#vision-language models#street view imagery#human‑AI agreement
astro-ph.CO2026

Constraining the Potential Index of the Early Dark Energy Model Using Cosmic Birefringence from Planck and ACT

Kedi Zhang, Lu Yin

Cosmic birefringence and the Hubble tension represent compelling challenges to the standard CDM model. The early dark energy (EDE) model with potentials $V(ϕ) \propto [1-\cos(…

cs.LG2025

Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Lu Yin, You Wu, Zhenyu Zhang +10

Large Language Models (LLMs), renowned for their remarkable performance across diverse domains, present a challenge when it comes to practical deployment due to their colossal mode…

astro-ph.CO2017

Constraints on running vacuum model with and

Chao-Qiang Geng, Chung-Chi Lee, Lu Yin

We examine the running vacuum model with , where is the model parameter and is the cosmological constant. From the data of the cosmic microwave…

cs.LG2025

LOST: Low-rank and Sparse Pre-training for Large Language Models

Jiaxi Li, Lu Yin, Li Shen +6

While large language models (LLMs) have achieved remarkable performance across a wide range of tasks, their massive scale incurs prohibitive computational and memory costs for pre-…

cs.CV2025

CLIMB-3D: Continual Learning for Imbalanced 3D Instance Segmentation

Vishal Thengane, Jean Lahoud, Hisham Cholakkal +4

While 3D instance segmentation (3DIS) has advanced significantly, most existing methods assume that all object classes are known in advance and uniformly distributed. However, this…

cs.CV2026

SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation

Vishal Thengane, Zhaochong An, Tianjin Huang +5

Incremental Few-Shot (IFS) segmentation aims to learn new categories over time from only a few annotations. Although widely studied in 2D, it remains underexplored for 3D point clo…

cs.LG2026

W2T: LoRA Weights Already Know What They Can Do

Xiaolong Han, Ferrante Neri, Zijian Jiang +4

Each LoRA checkpoint compactly stores task-specific updates in low-rank weight matrices, offering an efficient way to adapt large language models to new tasks and domains. In princ…

cs.LG2025

Outlier-weighed Layerwise Sampling for LLM Fine-tuning

Pengxiang Li, Lu Yin, Xiaowei Gao +1

The rapid advancements in Large Language Models (LLMs) have revolutionized various natural language processing tasks. However, the substantial size of LLMs presents significant cha…

astro-ph.CO2020

Reducing the tension with generalized Proca theory

Antonio De Felice, Chao-Qiang Geng, Masroor C. Pookkillath +1

We investigate the cosmological viability of the generalized proca theory. We first implement the background and linear perturbation equations of motion in the Boltzmann code and t…

cs.LG2026

A Survey of Weight Space Learning: Understanding, Representation, and Generation

Xiaolong Han, Zehong Wang, Bo Zhao +8

Neural network weights are typically viewed as the end product of training, while most deep learning research focuses on data, features, and architectures. However, recent advances…

cs.CL2026

Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?

Pengxiang Li, Dilxat Muhtar, Tianlong Chen +2

Diffusion Language Models (DLMs) are often advertised as enabling parallel token generation, yet practical fast DLMs frequently converge to left-to-right, autoregressive (AR)-like…

astro-ph.CO2015

A closer look at interacting dark energy with statefinder hierarchy and growth rate of structure

Jing-Lei Cui, Lu Yin, Ling-Feng Wang +2

We investigate the interacting dark energy models by using the diagnostics of statefinder hierarchy and growth rate of structure. We wish to explore the deviations from CDM and…

cs.CV2024

Are Sparse Neural Networks Better Hard Sample Learners?

Qiao Xiao, Boqian Wu, Lu Yin +4

While deep learning has demonstrated impressive progress, it remains a daunting challenge to learn from hard samples as these samples are usually noisy and intricate. These hard sa…

cs.LG2025

OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework

Jiaxi Li, Lu Yin, Xilu Wang

The integration of Large Language Models (LLMs) into autonomous driving systems offers promising enhancements in environmental understanding and decision-making. However, the subst…

cs.CV2024

MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization

Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma +5

Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facil…

cs.CL2026

Progressive Residual Warmup for Language Model Pretraining

Tianhao Chen, Xin Xu, Lu Yin +4

Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated…

cs.CL2025

AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs

Di He, Songjun Tu, Ajay Jaiswal +4

Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overl…

cs.GR2025

Shape-aware Inertial Poser: Motion Tracking for Humans with Diverse Shapes Using Sparse Inertial Sensors

Lu Yin, Ziying Shi, Yinghao Wu +3

Human motion capture with sparse inertial sensors has gained significant attention recently. However, existing methods almost exclusively rely on a template adult body shape to mod…

cs.CL2025

GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching

Guinan Su, Li Shen, Lu Yin +3

Large language models (LLMs) have shown remarkable capabilities in language understanding and generation. However, such impressive capability typically comes with a substantial mod…

hep-ph2020

Multicomponent Dark Matter in the Light of CALET and DAMPE

Chao-Qiang Geng, Da Huang, Lu Yin

In the light of the latest measurements on the total flux by CALET and DAMPE experiments, we revisit the multicomponent leptonically decaying dark matter (DM) explanati…