Publications (118)
A Dark Energy model from Generalized Proca Theory
Chao-Qiang Geng, Yan-Ting Hsu, Jhih-Rong Lu +1
We consider a specific dark energy model, which only includes the Lagrangian up to the cubic order in terms of the vector field self-interactions in the generalized Proca theory. W…
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
Hanqi Yan, Xiangxiang Cui, Lu Yin +4
The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing align…
Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing
Zihao Wang, Enneng Yang, Lu Yin +2
Model merging leverages multiple finetuned expert models to construct a multi-task model with low cost, and is gaining increasing attention. However, as a growing number of finetun…
TODO: Enhancing LLM Alignment with Ternary Preferences
Yuxiang Guo, Lu Yin, Bo Jiang +1
Aligning large language models (LLMs) with human intent is critical for enhancing their performance across a variety of tasks. Standard alignment techniques, such as Direct Prefere…
MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic Disturbances
Yunzhe Shao, Xinyu Yi, Lu Yin +3
This paper proposes a novel method called MagShield, designed to address the issue of magnetic interference in sparse inertial motion capture (MoCap) systems. Existing Inertial Mea…
Does Hubble Tension Signal a Breakdown in FLRW Cosmology?
Chethan Krishnan, Roya Mohayaee, Eoin à Colgáin +2
The tension between early and late Universe probes of the Hubble constant has motivated various new FLRW cosmologies. Here, we reanalyse the Hubble tension with a recent age of the…
Modified Cosmology Models from Thermodynamical Approach
Chao-Qiang Geng, Yan-Ting Hsu, Jhih-Rong Lu +1
We apply the first law of thermodynamics to the apparent horizon of the universe with the power-law corrected and non-extensive Tsallis entropies rather than the Bekenstein-Hawking…
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
Zhenyu Zhang, Ajay Jaiswal, Lu Yin +4
Training Large Language Models (LLMs) is memory-intensive due to the large number of parameters and associated optimization states. GaLore, a recent method, reduces memory usage by…
On Larger Values in the CMB Dipole Direction
Orlando Luongo, Marco Muccino, Eoin à Colgáin +2
On the assumption that quasars (QSO) and gamma-ray bursts (GRB) represent \textit{standardisable candles}, we provide evidence that the Hubble constant adopts larger values i…
Observational Constraints on the Cosmology with Holographic Dark Fluid
Da Huang, Bum-Hoon Lee, Gansukh Tumurtushaa +2
We consider the holographic Friedman-Robertson-Walker (hFRW) universe on the 4-dimensional membrane embedded in the 5-dimensional bulk spacetime and fit the parameters with the obs…
SMILE: a universal tool for modulated-enhanced localization microscopy to achieve minimal three-dimensional resolution
Hongfei Zhu, Yile Sun, Xinxun Yang +8
Modulation-enhanced localization microscopy (MELM) has demonstrated significant improvements in both lateral and axial localization precision compared to conventional single-molecu…
Semantic-Based Few-Shot Learning by Interactive Psychometric Testing
Lu Yin, Vlado Menkovski, Yulong Pei +1
Few-shot classification tasks aim to classify images in query sets based on only a few labeled examples in support sets. Most studies usually assume that each image in a task has a…
DymSLAM:4D Dynamic Scene Reconstruction Based on Geometrical Motion Segmentation
Chenjie Wang, Bin Luo, Yun Zhang +6
Most SLAM algorithms are based on the assumption that the scene is static. However, in practice, most scenes are dynamic which usually contains moving objects, these methods are no…
Constraints on a special running vacuum model
Chao-Qiang Geng, Chung-Chi Lee, Lu Yin
We study a special running vacuum model (RVM) with , where , and are the model parameters and is the Hubble one. This RVM has…
Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution
Qiao Xiao, Alan Ansell, Boqian Wu +4
Sparse large language models (LLMs) offer an attractive direction toward efficient deployment, but adapting them to downstream tasks remains challenging. The central difficulty is…
Is Cosmic Birefringence model-dependent?
Lu Yin, Joby Kochappan, Tuhin Ghosh +1
Exciting clues to isotropic cosmic birefringence have recently been detected in the cross-power spectra of the polarization data of the cosmic microwave background (CMB). Earl…
Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
Shiwei Liu, Lu Yin, Decebal Constantin Mocanu +1
In this paper, we introduce a new perspective on training deep neural networks capable of state-of-the-art performance without the need for the expensive over-parameterization by p…
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
Xin Xu, Yan Xu, Tianhao Chen +9
Existing approaches to mathematical reasoning with large language models (LLMs) rely on Chain-of-Thought (CoT) for generalizability or Tool-Integrated Reasoning (TIR) for precise c…
CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization
Zheyan Qu, Lu Yin, Zitong Yu +2
Large language models (LLMs) have demonstrated astonishing capabilities in natural language processing (NLP) tasks, sparking interest in their application to professional domains w…
The Curse of Depth in Large Language Models
Wenfang Sun, Xinyuan Song, Pengxiang Li +3
In this paper, we introduce the Curse of Depth, a concept that highlights, explains, and addresses the recent observation in modern Large Language Models (LLMs) where nearly half o…
Design and Performance Analysis of Multi-scale NOMA for 5G Positioning
Lu Yin, Jiameng Cao, Zhongliang Deng +4
This paper presents a feasibility study for a novel positioning-communication integrated signal called Multi-Scale Non-Orthogonal Multiple Access (MS-NOMA) for 5G positioning. One…
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
Lu Yin, Ajay Jaiswal, Shiwei Liu +2
We present Junk DNA Hypothesis by adopting a novel task-centric angle for the pre-trained weights of large language models (LLMs). It has been believed that weights in LLMs contain…
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
Abhinav Bandari, Lu Yin, Cheng-Yu Hsieh +5
Network pruning has emerged as a potential solution to make LLMs cheaper to deploy. However, existing LLM pruning approaches universally rely on the C4 dataset as the calibration d…
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Di He, Songjun Tu, Keyu Wang +2
Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural…
WIMPs in Dilatonic Einstein Gauss-Bonnet Cosmology
Anirban Biswas, Arpan Kar, Bum-Hoon Lee +5
We use the Weakly Interacting Massive Particle (WIMP) thermal decoupling scenario to probe Cosmologies in dilatonic Einstein Gauss-Bonnet (dEGB) gravity, where the Gauss-Bonnet ter…
Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
Xin Xu, Tianhao Chen, Fan Zhang +11
While slow-thinking large language models (LLMs) exhibit reflection-like reasoning, commonly referred to as the "aha moment:, their ability to generate informative critiques and re…
E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation
Boqian Wu, Qiao Xiao, Shiwei Liu +5
Deep neural networks have evolved as the leading approach in 3D medical image segmentation due to their outstanding performance. However, the ever-increasing model size and computa…
High Performance Printed AgO-Zn Rechargeable Battery for Flexible Electronics
Lu Yin, Jonathan Scharf, Jessica Ma +8
The rise of flexible electronics calls for cost-effective and scalable batteries with good mechanical and electrochemical performance. In this work, we developed printable, polymer…
You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets
Tianjin Huang, Tianlong Chen, Meng Fang +8
Recent works have impressively demonstrated that there exists a subnetwork in randomly initialized convolutional neural networks (CNNs) that can match the performance of the fully…
Robust Active Learning (RoAL): Countering Dynamic Adversaries in Active Learning with Elastic Weight Consolidation
Ricky Maulana Fajri, Yulong Pei, Lu Yin +1
Despite significant advancements in active learning and adversarial attacks, the intersection of these two fields remains underexplored, particularly in developing robust active le…
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
Xinchen Han, Hossam Afifi, Michel Marot +2
Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency without proportional performance g…
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
Zihang Liu, Tianyu Pang, Oleg Balabanov +5
Recent studies have shown that supervised fine-tuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabilities. However, full fine-tuning (Full FT…
Inevitable manifestation of wiggles in the expansion of the late Universe
Ozgur Akarsu, Eoin O. Colgain, Emre Ozulker +2
Using the fact that the comoving angular diameter distance to the last scattering surface is strictly constrained almost model independently, we show that, for any model agreeing w…
Hints of FLRW Breakdown from Supernovae
Chethan Krishnan, Roya Mohayaee, Eoin à Colgáin +2
A 10\% difference in the scale for the Hubble parameter constitutes a clear problem for cosmology. Here, considering angular distribution of Type Ia supernovae (SN) within the Pant…
TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
Shaojie Zhuang, Lu Yin, Guangshun Wei +3
Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing approaches rely on task-specific…
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
Ajay Jaiswal, Yifan Wang, Lu Yin +6
Large Language Models' (LLMs) weight matrices can often be expressed in low-rank form with potential to relax memory and compute resource requirements. Unlike prior efforts that fo…
Are Large Kernels Better Teachers than Transformers for ConvNets?
Tianjin Huang, Lu Yin, Zhenyu Zhang +5
This paper reveals a new appeal of the recently emerged large-kernel Convolutional Neural Networks (ConvNets): as the teacher in Knowledge Distillation (KD) for small-kernel ConvNe…
Into the Unknown: Applying Inductive Spatial-Semantic Location Embeddings for Predicting Individuals' Mobility Beyond Visited Places
Xinglei Wang, Tao Cheng, Stephen Law +6
Predicting individuals' next locations is a core task in human mobility modelling, with wide-ranging implications for urban planning, transportation, public policy and personalised…
Pushing the Limits of Sparsity: A Bag of Tricks for Extreme Pruning
Andy Li, Aiden Durrant, Milan Markovic +5
Pruning of deep neural networks has been an effective technique for reducing model size while preserving most of the performance of dense networks, crucial for deploying models on…
Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
Zhaorui Meng, Lu Yin, Yangqing Hou +3
Sparse Inertial Measurement Units (IMUs) based human motion capture has gained significant momentum, driven by the adaptation of fundamental AI tools such as recurrent neural netwo…
Dynamic Data Pruning for Automatic Speech Recognition
Qiao Xiao, Pingchuan Ma, Adriana Fernandez-Lopez +7
The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitivel…
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework
Zhuo Zhi, Chen Feng, Adam Daneshmend +6
Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic fr…
Reducing the Tension with Exponential Acoustic Dark Energy
Lu Yin
The Hubble tension arises from different observations between the late-time and early Universe. We explore a new model with dark fluid, called the exponential Acoustic Dark Energy…
Constraints on Cosmic Birefringence from SPIDER, Planck, and ACT observations
Lu Yin, Shuhang Xiong, Joby Kochappan +2
The Early Dark Energy (EDE) model has been proposed as a candidate mechanism to generate cosmic birefringence through a Chern-Simons coupling between a dynamical scalar field and t…
Observational evidence for Early Dark Energy as a unified explanation for Cosmic Birefringence and the Hubble tension
Joby Kochappan, Lu Yin, Bum-Hoon Lee +1
We test the =3 Ultralight Axion-like model of Early Dark Energy (EDE) with the observationsof the mode of the cosmic microwave background (CMB) radiation, and local expansi…
Does Gauss-Bonnet Inflationary Gravitational Waves satisfy the Pulsar Timing Arrays observations?
Lu Yin
The observations from pulsar timing arrays (PTAs), led by the North American Nanohertz Observatory for Gravitational Waves (NANOGrav), have provided opportunities to constrain prim…
Linear-Time Self Attention with Codeword Histogram for Efficient Recommendation
Yongji Wu, Defu Lian, Neil Zhenqiang Gong +4
Self-attention has become increasingly popular in a variety of sequence modeling tasks from natural language processing to recommendation, due to its effectiveness. However, self-a…
Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
Shiwei Liu, Tianlong Chen, Xiaohan Chen +7
Works on lottery ticket hypothesis (LTH) and single-shot network pruning (SNIP) have raised a lot of attention currently on post-training pruning (iterative magnitude pruning), and…
Light quark energy loss in the flavor-dependent systems from holography
Le Zhang, Lu Yin, Guo-Dong Zhou +2
Using the holographic model of finite-endpoint-momentum shooting string approach, we study the instantaneous energy loss of light quarks for the flavor-dependent systems with $N_f…
MedFM-Robust: Benchmarking Robustness of Medical Foundation Models
Xiangxiang Cui, Tianjin Huang, Yifang Wang +2
Medical foundation models have achieved remarkable clinical performance, yet their robustness under real-world perturbations remains underexplored. We present a robustness benchmar…
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
Zihan Wang, Rui Pan, Jiarui Yao +7
We propose Chain-of-Experts (CoE), a new Mixture-of-Experts (MoE) architecture that introduces sequential expert communication within each layer. Unlike traditional MoE models, whe…
Late-time Cosmology without Dark Sector but with Closed String Massless Sector
Hocheol Lee, Jeong-Hyuck Park, Liliana Velasco-Sevilla +1
We explore the possibility of solving the dark energy and the coincidence problems by postulating the massless sector of closed strings. This sector constitutes the gravitational m…
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Pengxiang Li, Lu Yin, Shiwei Liu
Large Language Models (LLMs) have achieved remarkable success, yet recent findings reveal that their deeper layers often contribute minimally and can be pruned without affecting ov…
SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment
Ziyang Chen, Zhenxuan Huang, Yile Wang +3
Traditional sentence embedding methods employ token-level contrastive learning on non-generative pre-trained models. Recently, there have emerged embedding methods based on generat…
Do high redshift QSOs and GRBs corroborate JWST?
Eoin à Colgáin, M. M. Sheikh-Jabbari, Lu Yin
The James Webb Space Telescope (JWST) is reporting massive high redshift galaxies that appear challenging from the CDM perspective. Interpreted as a cosmological problem, this…
A Structural-Clustering Based Active Learning for Graph Neural Networks
Ricky Maulana Fajri, Yulong Pei, Lu Yin +1
In active learning for graph-structured data, Graph Neural Networks (GNNs) have shown effectiveness. However, a common challenge in these applications is the underutilization of cr…
Hierarchical Semantic Segmentation using Psychometric Learning
Lu Yin, Vlado Menkovski, Shiwei Liu +1
Assigning meaning to parts of image data is the goal of semantic image segmentation. Machine learning methods, specifically supervised learning is commonly used in a variety of tas…
CPR: Chained Perceptual Refinement for Coarse-to-Fine Medical Image Classification
Si-Yuan Lu, Hanruo Zhu, Ziquan Zhu +6
High resolution medical images contain fine grained, spatially sparse cues that are critical for diagnosis, yet preserving full resolution incurs substantial computational and memo…
How much has DESI dark energy evolved since DR1?
Eoin à Colgáin, Saeed Pourojaghi, M. M. Sheikh-Jabbari +1
DESI has reported a dynamical dark energy (DE) signal based on the CDM model that is in conflict with Hubble tension. Recalling that the combination of DESI DR1 BAO and DR…
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
Ajay Jaiswal, Bodun Hu, Lu Yin +4
Autoregressive Large Language Models (e.g., LLaMa, GPTs) are omnipresent achieving remarkable success in language understanding and generation. However, such impressive capability…
Gravitational waves from the vacuum decay with LISA
Bum-Hoon Lee, Wonwoo Lee, Dong-han Yeom +1
We investigate the gravitational wave spectrum resulted from the cosmological first-order phase transition. We compare two models; one is a scalar field model without gravitation,…
Rethinking Lifelong Sequential Recommendation with Incremental Multi-Interest Attention
Yongji Wu, Lu Yin, Defu Lian +4
Sequential recommendation plays an increasingly important role in many e-commerce services such as display advertisement and online shopping. With the rapid development of these se…
The CosmoVerse White Paper: Addressing observational tensions in cosmology with systematics and fundamental physics
Eleonora Di Valentino, Jackson Levi Said, Adam Riess +533
The standard model of cosmology has provided a good phenomenological description of a wide range of observations both at astrophysical and cosmological scales for several decades.…
Investigating Degradation Modes in Zn-AgO Aqueous Batteries with X-ray Micro Computed Tomography
Jonathan Scharf, Lu Yin, Christopher Redquest +7
To meet growing energy demands, degradation mechanisms of energy storage devices must be better understood. As a non-destructive tool, X-ray Computed Tomography (CT) has been incre…
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Keyu Wang, Tian Lyu, Guinan Su +4
Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing methods demonstrate strong performance reten…
Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
Xinyuan Song, Keyu Wang, PengXiang Li +2
Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance…
Knowledge Elicitation using Deep Metric Learning and Psychometric Testing
Lu Yin, Vlado Menkovski, Mykola Pechenizkiy
Knowledge present in a domain is well expressed as relationships between corresponding concepts. For example, in zoology, animal species form complex hierarchies; in genomics, the…
Can dark energy be dynamical?
Eoin à Colgáin, M. M. Sheikh-Jabbari, Lu Yin
We highlight shortcomings of the dynamical dark energy (DDE) paradigm. For parametric models with equation of state (EOS), for a given function of redshift…
Aspect-Based Few-Shot Learning
Tim van Engeland, Lu Yin, Vlado Menkovski
We generalize the formulation of few-shot learning by introducing the concept of an aspect. In the traditional formulation of few-shot learning, there is an underlying assumption t…
Joint constraints on cosmic birefringence and early dark energy from ACT, Planck, DESI, and PantheonPlus
Lu Yin, Guo-Hong Du, Tian-Nuo Li +1
With the increasing number of high-precision astronomical observations, physical quantities that were previously inaccessible to accurate calculations, such as cosmic birefringence…
Cosmological constraints on CDM models with time-varying fine structure constant
Jin-Jun Zhang, Lu Yin, Chao-Qiang Geng
We study the CDM models with being a function of the time-varying fine structure constant . We give a close look at the constraints on two specific CDM…
Multimodal Contrastive Learning of Urban Space Representations from POI Data
Xinglei Wang, Tao Cheng, Stephen Law +3
Existing methods for learning urban space representations from Point-of-Interest (POI) data face several limitations, including issues with geographical delineation, inadequate spa…
Stochastic Gravitational Waves from Inflaton Decays
Da Huang, Lu Yin
Due to the universality of gravitational interactions, it is generally expected that a stochastic gravitational wave (GW) background could form during the reheating period when the…
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
Jiaxi Li, Lu Yin, Li Shen +5
Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck for widespread adoption. Lo…
Sebra: Debiasing Through Self-Guided Bias Ranking
Adarsh Kappiyath, Abhra Chaudhuri, Ajay Jaiswal +4
Ranking samples by fine-grained estimates of spuriosity (the degree to which spurious cues are present) has recently been shown to significantly benefit bias mitigation, over the t…
Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Models
Mingyu Cao, Alvaro H. C. Correia, Christos Louizos +2
Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, repeatedly deciding which positions to commit at each step. Standard decoding follows a g…
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
Tianhao Chen, Xin Xu, Zijing Liu +12
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretra…
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin +2
This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viabi…
Running vacuum model in non-flat universe
Chao-Qiang Geng, Yan-Ting Hsu, Lu Yin +1
We investigate observational constraints on the running vacuum model (RVM) of in the spatially curved universe, where is the model parameter, cor…
AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
Zahratu Shabrina, Muhammad Asa, Jin Rui +2
The paper evaluates how state‑of‑the‑art vision‑language models (e.g., GPT‑4o, Claude 3.5 Sonnet, Gemini 2.0 Flash) predict building typology attributes from Google Street View ima…
Constraining the Potential Index of the Early Dark Energy Model Using Cosmic Birefringence from Planck and ACT
Kedi Zhang, Lu Yin
Cosmic birefringence and the Hubble tension represent compelling challenges to the standard CDM model. The early dark energy (EDE) model with potentials $V(Ï) \propto [1-\cos(…
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Lu Yin, You Wu, Zhenyu Zhang +10
Large Language Models (LLMs), renowned for their remarkable performance across diverse domains, present a challenge when it comes to practical deployment due to their colossal mode…
Constraints on running vacuum model with and
Chao-Qiang Geng, Chung-Chi Lee, Lu Yin
We examine the running vacuum model with , where is the model parameter and is the cosmological constant. From the data of the cosmic microwave…
LOST: Low-rank and Sparse Pre-training for Large Language Models
Jiaxi Li, Lu Yin, Li Shen +6
While large language models (LLMs) have achieved remarkable performance across a wide range of tasks, their massive scale incurs prohibitive computational and memory costs for pre-…
CLIMB-3D: Continual Learning for Imbalanced 3D Instance Segmentation
Vishal Thengane, Jean Lahoud, Hisham Cholakkal +4
While 3D instance segmentation (3DIS) has advanced significantly, most existing methods assume that all object classes are known in advance and uniformly distributed. However, this…
SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
Vishal Thengane, Zhaochong An, Tianjin Huang +5
Incremental Few-Shot (IFS) segmentation aims to learn new categories over time from only a few annotations. Although widely studied in 2D, it remains underexplored for 3D point clo…
W2T: LoRA Weights Already Know What They Can Do
Xiaolong Han, Ferrante Neri, Zijian Jiang +4
Each LoRA checkpoint compactly stores task-specific updates in low-rank weight matrices, offering an efficient way to adapt large language models to new tasks and domains. In princ…
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
Pengxiang Li, Lu Yin, Xiaowei Gao +1
The rapid advancements in Large Language Models (LLMs) have revolutionized various natural language processing tasks. However, the substantial size of LLMs presents significant cha…
Reducing the tension with generalized Proca theory
Antonio De Felice, Chao-Qiang Geng, Masroor C. Pookkillath +1
We investigate the cosmological viability of the generalized proca theory. We first implement the background and linear perturbation equations of motion in the Boltzmann code and t…
A Survey of Weight Space Learning: Understanding, Representation, and Generation
Xiaolong Han, Zehong Wang, Bo Zhao +8
Neural network weights are typically viewed as the end product of training, while most deep learning research focuses on data, features, and architectures. However, recent advances…
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
Pengxiang Li, Dilxat Muhtar, Tianlong Chen +2
Diffusion Language Models (DLMs) are often advertised as enabling parallel token generation, yet practical fast DLMs frequently converge to left-to-right, autoregressive (AR)-like…
A closer look at interacting dark energy with statefinder hierarchy and growth rate of structure
Jing-Lei Cui, Lu Yin, Ling-Feng Wang +2
We investigate the interacting dark energy models by using the diagnostics of statefinder hierarchy and growth rate of structure. We wish to explore the deviations from CDM and…
Are Sparse Neural Networks Better Hard Sample Learners?
Qiao Xiao, Boqian Wu, Lu Yin +4
While deep learning has demonstrated impressive progress, it remains a daunting challenge to learn from hard samples as these samples are usually noisy and intricate. These hard sa…
OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework
Jiaxi Li, Lu Yin, Xilu Wang
The integration of Large Language Models (LLMs) into autonomous driving systems offers promising enhancements in environmental understanding and decision-making. However, the subst…
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma +5
Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facil…
Progressive Residual Warmup for Language Model Pretraining
Tianhao Chen, Xin Xu, Lu Yin +4
Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated…
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Di He, Songjun Tu, Ajay Jaiswal +4
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overl…
Shape-aware Inertial Poser: Motion Tracking for Humans with Diverse Shapes Using Sparse Inertial Sensors
Lu Yin, Ziying Shi, Yinghao Wu +3
Human motion capture with sparse inertial sensors has gained significant attention recently. However, existing methods almost exclusively rely on a template adult body shape to mod…
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
Guinan Su, Li Shen, Lu Yin +3
Large language models (LLMs) have shown remarkable capabilities in language understanding and generation. However, such impressive capability typically comes with a substantial mod…
Multicomponent Dark Matter in the Light of CALET and DAMPE
Chao-Qiang Geng, Da Huang, Lu Yin
In the light of the latest measurements on the total flux by CALET and DAMPE experiments, we revisit the multicomponent leptonically decaying dark matter (DM) explanati…