#gpu acceleration
22 papers match
A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes
Alper Sahistan, Haichao Miao, Zhimin Li +3
The paper introduces a stochastic volume rendering framework that efficiently renders time‑varying implicit neural volumes by reducing neural network queries with delta tracking, r…
Hash Chemistry: Minimal Models for Evolutionary Growth of Complexity
Ilya Horiguchi, Hiroki Sayama
The paper reviews the Hash Chemistry framework of minimal evolutionary models that use hash functions to assign fitness scores, and extends the Structural Cellular Hash Chemistry m…
Advancing Awkward Arrays for High-Performance CPU and GPU Processing
Ianna Osborne, Manasvi Goyal
The paper describes recent enhancements to the Awkward Array Python library that enable efficient GPU processing of nested, variable‑length data using CUDA and NVIDIA's core comput…
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
Yuxin Chen, Liang Luo, Buyun Zhang +44
The paper introduces ROCS, a request-oriented compute sharing framework that restructures recommendation inference to evaluate shared request features once per request rather than…
HORSES3D-GPU: A high-order discontinuous Galerkin solver for multi-GPU systems
Gerasimos Ntoukas, Gonzalo Rubio, Abbas Ballout +13
The paper describes the GPU-accelerated version of the open-source high-order discontinuous Galerkin CFD solver HORSES3D, showing its performance and scalability on multi‑GPU syste…
InferScale: GPU-Native KV Injection for Personalized LLM Serving
Peter Li, Prashant Pandey
The paper introduces InferScale, a GPU-native system that precomputes and caches key‑value (KV) representations of personalized memory facts for large language models, allowing dir…
Pictura: Perspective-View Self-Play at Scale for Driving
Yuan Yin, Elias Ramzi, Marc Lafon +8
The paper presents Pictura, a GPU‑accelerated multi‑agent driving simulator that renders each vehicle's egocentric camera view, enabling large‑scale self‑play training of driving p…
Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging
Christopher Hahne
The paper proposes Quasi‑SVD, a differentiable matrix factorisation that enforces orthogonality on a single Lie‑parameterized factor, enabling fully parallel GPU computation and ac…
Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study
Lukas Stepanek
The paper studies how the alignment of routed tokens into expert blocks determines the exact packed quantized matrix multiplication performed during Mixture‑of‑Experts inference, d…
phase2: Full-State Vector Simulation of Quantum Time Evolution at Scale
Marek Miller, Jakob Günther, Freek Witteveen +5
phase2 is a full-state-vector quantum simulator that efficiently handles many-qubit Pauli rotation circuits on distributed CPU/GPU clusters, achieving large speedups and scaling to…
Fast two-dimensional tensor-network contraction via subspace iteration
Yining Zhang, Philippe Corboz
The paper proposes a subspace-iteration CTMRG method that replaces large SVDs with smaller ones using QR-based projectors, greatly speeding up iPEPS tensor-network contractions and…
Differentiable Routability-Driven Package Floorplanning with Pin Assignment
Yiqi Huang, Zepeng Li, Zhen Zhuang +3
The paper introduces a differentiable optimization framework for package floorplanning and pin assignment that directly models chip orientations and fan‑out congestion, achieving f…
Compression of 3D Gaussian Splatting Data Using GPU-friendly Graphics Texture Coding
Amir Said, Randall Rauwendaal
The paper proposes GPU-friendly texture compression methods to reduce the memory needed for spherical harmonic color coefficients in 3D Gaussian Splatting, using BC1/BC7 formats an…
Realizable N:M Sparse Transformer Inference via Search-Kernel Co-Design
Yiming Liu, Wenqi Lou, Zhiguang Wang +4
The paper presents a co-designed hardware and software framework that enables fast inference of Vision Transformers by applying N:M structured sparsity with a specialized CUDA kern…
3DPipe: A Pipelined GPU Framework for Scalable Generalized Spatial Join over Polyhedral Objects
Lyuheng Yuan, Da Yan, Saugat Adhikari +2
The paper introduces 3DPipe, a GPU-based framework that pipelines filtering and refinement to efficiently perform spatial joins on large 3D polyhedral datasets, using multi-level p…
DRIFT: Direct Reduced Fourier Transforms for Distributed Spectral Neural Operators
Sana Taghipour Anvari, David Kaeli
The paper introduces DRIFT, a GPU implementation that computes only the needed frequency modes for Fourier Neural Operators using a Distributed Truncated Spectral Transform, dramat…
GPU-Accelerated Optimisation of Symmetric Bidirectional Ultrawideband Coherent Transmission Under Launch Power Constraints
Mindaugas JarmoloviÄius, Eric Sillekens, Polina Bayvel +1
The paper presents a GPU‑accelerated optimisation framework for bidirectional ultrawideband coherent transmission under total fibre launch‑power constraints, demonstrating up to a…
CayleyR: Solving the TopSpin puzzle via cycle intersection
Yuri Baramykov
The paper introduces cayleyR, an R package that solves permutation puzzles such as TopSpin by using an iterative bidirectional search to find intersecting cycles in Cayley graphs,…
JAX-FEM-ANISO: Differentiable GPU-Accelerated Finite Element Framework for Inverse Identification of Finite-Strain Anisotropic Plasticity
Deepak Sharma, Itzel Salgado, Lu Huang +2
The paper introduces JAX-FEM-ANISO, a fully differentiable, GPU‑accelerated finite element framework that enables fast forward simulations and inverse identification of finite‑stra…
WarpMPC: Large-Batch MPC on GPU via ADMM with Unrolled Factorization
Henrik Hose, Se Hwan Jeon, Charles Khazoom +2
The paper introduces WarpMPC, a GPU‑accelerated toolbox that speeds up large‑batch model predictive control by unrolling sparse LDLᵀ factorizations within an ADMM solver, achieving…
DeepRT Engine: A Unified GPU-Parallel Ray-Tracing Framework with Hybrid SBR-IM Path Search for 6G Digital Twin Channel
Tao Wu, Li Yu, Yuxiang Zhang +3
The paper introduces DeepRT Engine, a GPU-parallel ray‑tracing framework that speeds up the creation of real‑time digital twin wireless channels by using a three‑stage pipeline wit…
CR-Solver: GPU-Accelerated Kinematics Solver for Tendon-driven Continuum Robots
Heqing Yang, Yang Yi, Linqing Zhong +2
The paper introduces CR-Solver, a GPU‑accelerated optimization framework that solves inverse kinematics, path following, and trajectory planning for tendon‑driven continuum robots…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.