#gpu acceleration

try —

22 papers match

cs.GR2026

A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes

Alper Sahistan, Haichao Miao, Zhimin Li +3

The paper introduces a stochastic volume rendering framework that efficiently renders time‑varying implicit neural volumes by reducing neural network queries with delta tracking, r…

#volume rendering#implicit neural representations#time-varying data#stochastic rendering
q-bio.PE2026

Hash Chemistry: Minimal Models for Evolutionary Growth of Complexity

Ilya Horiguchi, Hiroki Sayama

The paper reviews the Hash Chemistry framework of minimal evolutionary models that use hash functions to assign fitness scores, and extends the Structural Cellular Hash Chemistry m…

#open-ended evolution#minimal evolutionary models#hash chemistry#spatial locality
cs.SE2026

Advancing Awkward Arrays for High-Performance CPU and GPU Processing

Ianna Osborne, Manasvi Goyal

The paper describes recent enhancements to the Awkward Array Python library that enable efficient GPU processing of nested, variable‑length data using CUDA and NVIDIA's core comput…

#gpu acceleration#nested arrays#ragged data structures#cuda backend
cs.LG2026

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

Yuxin Chen, Liang Luo, Buyun Zhang +44

The paper introduces ROCS, a request-oriented compute sharing framework that restructures recommendation inference to evaluate shared request features once per request rather than…

#recommendation#inference optimization#compute sharing#large-scale systems
math.NA2026

HORSES3D-GPU: A high-order discontinuous Galerkin solver for multi-GPU systems

Gerasimos Ntoukas, Gonzalo Rubio, Abbas Ballout +13

The paper describes the GPU-accelerated version of the open-source high-order discontinuous Galerkin CFD solver HORSES3D, showing its performance and scalability on multi‑GPU syste…

#high-order methods#discontinuous galerkin#gpu acceleration#large-scale simulation
cs.DC2026

InferScale: GPU-Native KV Injection for Personalized LLM Serving

Peter Li, Prashant Pandey

The paper introduces InferScale, a GPU-native system that precomputes and caches key‑value (KV) representations of personalized memory facts for large language models, allowing dir…

#large language models#kv cache#personalized memory#inference optimization
cs.CV2026

Pictura: Perspective-View Self-Play at Scale for Driving

Yuan Yin, Elias Ramzi, Marc Lafon +8

The paper presents Pictura, a GPU‑accelerated multi‑agent driving simulator that renders each vehicle's egocentric camera view, enabling large‑scale self‑play training of driving p…

#self-play#driving simulation#egocentric perception#reinforcement learning
cs.CV2026

Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

Christopher Hahne

The paper proposes Quasi‑SVD, a differentiable matrix factorisation that enforces orthogonality on a single Lie‑parameterized factor, enabling fully parallel GPU computation and ac…

#matrix factorisation#real-time imaging#gpu acceleration#lie groups
cs.DC2026

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study

Lukas Stepanek

The paper studies how the alignment of routed tokens into expert blocks determines the exact packed quantized matrix multiplication performed during Mixture‑of‑Experts inference, d…

#mixture-of-experts#quantized inference#routing alignment#packed arithmetic
quant-ph2026

phase2: Full-State Vector Simulation of Quantum Time Evolution at Scale

Marek Miller, Jakob Günther, Freek Witteveen +5

phase2 is a full-state-vector quantum simulator that efficiently handles many-qubit Pauli rotation circuits on distributed CPU/GPU clusters, achieving large speedups and scaling to…

#quantum simulation#state vector#pauli rotations#gpu acceleration
cond-mat.str-el2026

Fast two-dimensional tensor-network contraction via subspace iteration

Yining Zhang, Philippe Corboz

The paper proposes a subspace-iteration CTMRG method that replaces large SVDs with smaller ones using QR-based projectors, greatly speeding up iPEPS tensor-network contractions and…

#tensor networks#iPEPS#CTMRG#gpu acceleration
cs.AR2026

Differentiable Routability-Driven Package Floorplanning with Pin Assignment

Yiqi Huang, Zepeng Li, Zhen Zhuang +3

The paper introduces a differentiable optimization framework for package floorplanning and pin assignment that directly models chip orientations and fan‑out congestion, achieving f…

#package floorplanning#pin assignment#routability optimization#differentiable optimization
cs.CV2026

Compression of 3D Gaussian Splatting Data Using GPU-friendly Graphics Texture Coding

Amir Said, Randall Rauwendaal

The paper proposes GPU-friendly texture compression methods to reduce the memory needed for spherical harmonic color coefficients in 3D Gaussian Splatting, using BC1/BC7 formats an…

#3d gaussian splatting#texture compression#spherical harmonics#gpu acceleration
cs.AR2026

Realizable N:M Sparse Transformer Inference via Search-Kernel Co-Design

Yiming Liu, Wenqi Lou, Zhiguang Wang +4

The paper presents a co-designed hardware and software framework that enables fast inference of Vision Transformers by applying N:M structured sparsity with a specialized CUDA kern…

#vision transformers#n:m sparsity#hardware-software co-design#gpu acceleration
cs.DB2026

3DPipe: A Pipelined GPU Framework for Scalable Generalized Spatial Join over Polyhedral Objects

Lyuheng Yuan, Da Yan, Saugat Adhikari +2

The paper introduces 3DPipe, a GPU-based framework that pipelines filtering and refinement to efficiently perform spatial joins on large 3D polyhedral datasets, using multi-level p…

#spatial join#3d data#gpu acceleration#polyhedral objects
cs.DC2026

DRIFT: Direct Reduced Fourier Transforms for Distributed Spectral Neural Operators

Sana Taghipour Anvari, David Kaeli

The paper introduces DRIFT, a GPU implementation that computes only the needed frequency modes for Fourier Neural Operators using a Distributed Truncated Spectral Transform, dramat…

#fourier neural operators#distributed training#gpu acceleration#spectral methods
eess.SP2026

GPU-Accelerated Optimisation of Symmetric Bidirectional Ultrawideband Coherent Transmission Under Launch Power Constraints

Mindaugas Jarmolovičius, Eric Sillekens, Polina Bayvel +1

The paper presents a GPU‑accelerated optimisation framework for bidirectional ultrawideband coherent transmission under total fibre launch‑power constraints, demonstrating up to a…

#coherent transmission#bidirectional communication#ultrawideband#launch power optimization
cs.AI2026

CayleyR: Solving the TopSpin puzzle via cycle intersection

Yuri Baramykov

The paper introduces cayleyR, an R package that solves permutation puzzles such as TopSpin by using an iterative bidirectional search to find intersecting cycles in Cayley graphs,…

#permutation puzzles#cayley graphs#bidirectional search#topspin
cs.CE2026

JAX-FEM-ANISO: Differentiable GPU-Accelerated Finite Element Framework for Inverse Identification of Finite-Strain Anisotropic Plasticity

Deepak Sharma, Itzel Salgado, Lu Huang +2

The paper introduces JAX-FEM-ANISO, a fully differentiable, GPU‑accelerated finite element framework that enables fast forward simulations and inverse identification of finite‑stra…

#finite element method#anisotropic plasticity#automatic differentiation#gpu acceleration
cs.RO2026

WarpMPC: Large-Batch MPC on GPU via ADMM with Unrolled Factorization

Henrik Hose, Se Hwan Jeon, Charles Khazoom +2

The paper introduces WarpMPC, a GPU‑accelerated toolbox that speeds up large‑batch model predictive control by unrolling sparse LDLᵀ factorizations within an ADMM solver, achieving…

#gpu acceleration#model predictive control#large-batch optimization#admm
eess.SP2026

DeepRT Engine: A Unified GPU-Parallel Ray-Tracing Framework with Hybrid SBR-IM Path Search for 6G Digital Twin Channel

Tao Wu, Li Yu, Yuxiang Zhang +3

The paper introduces DeepRT Engine, a GPU-parallel ray‑tracing framework that speeds up the creation of real‑time digital twin wireless channels by using a three‑stage pipeline wit…

#digital twin#ray tracing#gpu acceleration#channel modeling
cs.RO2026

CR-Solver: GPU-Accelerated Kinematics Solver for Tendon-driven Continuum Robots

Heqing Yang, Yang Yi, Linqing Zhong +2

The paper introduces CR-Solver, a GPU‑accelerated optimization framework that solves inverse kinematics, path following, and trajectory planning for tendon‑driven continuum robots…

#continuum robots#tendon-driven actuation#kinematics solving#gpu acceleration

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.