• Cite Icon19
  • https://doi.org/10.1145/3309987Copy DOI Icon

SPIN

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recent GPUs enable Peer-to-Peer Direct Memory Access ( p 2 p ) from fast peripheral devices like NVMe SSDs to exclude the CPU from the data path between them for efficiency. Unfortunately, using p 2 p to access files is challenging because of the subtleties of low-level non-standard interfaces, which bypass the OS file I/O layers and may hurt system performance. Developers must possess intimate knowledge of low-level interfaces to manually handle the subtleties of data consistency and misaligned accesses. We present SPIN , which integrates p 2 p into the standard OS file I/O stack, dynamically activating p 2 p where appropriate, transparently to the user. It combines p 2 p with page cache accesses, re-enables read-ahead for sequential reads, all while maintaining standard POSIX FS consistency, portability across GPUs and SSDs, and compatibility with virtual block devices such as software RAID. We evaluate SPIN on NVIDIA and AMD GPUs using standard file I/O benchmarks, application traces, and end-to-end experiments. SPIN achieves significant performance speedups across a wide range of workloads, exceeding p 2 p throughput by up to an order of magnitude. It also boosts the performance of an aerial imagery rendering application by 2.6× by dynamically adapting to its input-dependent file access pattern, enables 3.3× higher throughput for a GPU-accelerated log server, and enables 29% faster execution for the highly optimized GPU-accelerated image collage with only 30 changed lines of code.

Similar Papers
  • Research Article
  • Citations2

DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor

  • Jun 30, 2025
  • ACM Transactions on Architecture and Code Optimization
  • Gia Bao Thieu +2
  • Conference Article
  • Citations49

DICE: Automatic Emulation of DMA Input Channels for Dynamic Firmware Analysis

  • May 01, 2021
  • Alejandro Mera +3
  • Research Article
  • Citations3

Architectural Support for Coherent Architecturally Visible Storage in Instruction Set Extensions

  • Jan 01, 2010
  • Infoscience (Ecole Polytechnique Fédérale de Lausanne)
  • Ties Jan Henderikus Kluter
  • Research Article
  • Citations20

New capabilities of the Monte Carlo dose engine ARCHER-RT: Clinical validation of the Varian TrueBeam machine for VMAT external beam radiotherapy.

  • Apr 13, 2020
  • Medical Physics
  • David P Adam +4
  • Research Article

整合SDRAM控制器、主僕式矽智財、與SoC匯流排分析儀之系統晶片設計架構

  • Jan 01, 2009
  • 吳俊慶
  • Research Article

DMA method for NAND flash based on PXA3xx processor

  • Oct 09, 2009
  • Journal of Computer Applications
  • Bin Shi +2
  • Book Chapter
  • Citations10

Evaluating Performance Portability of OpenMP for SNAP on NVIDIA, Intel, and AMD GPUs Using the Roofline Methodology

  • Jan 01, 2021
  • Neil A Mehta +4
  • Conference Article
  • Citations6

Design and Implementation of EDMA Controller for AI based DSP SoCs for Real- Time Multimedia Processing

  • Oct 07, 2020
  • Madhuri R A +5
  • Conference Article
  • Citations14

Poster

  • Oct 17, 2011
  • Patrick Stewin +2
  • Research Article
  • Citations24

DMA transfer method for wide-range speed and frequency measurement

  • Jan 01, 1993
  • IEEE Transactions on Instrumentation and Measurement
  • M Prokin
  • Research Article
  • Citations2

HiPeC — High Performance Cryptographic Service for Heterogeneous Network-on-Chip Systems

  • Jan 01, 2015
  • IFAC PapersOnLine
  • Hermann Seuschek +2
  • Conference Article
  • Citations14

A Processor-DMA-Based Memory Copy Hardware Accelerator

  • Jul 01, 2011
  • Wen Su +3
  • Research Article
  • Citations15

Optimizing Inter-Core Communications Under the LET Paradigm using DMA Engines

  • Jan 01, 2023
  • IEEE Transactions on Computers
  • Paolo Pazzaglia +3
  • Conference Article

Dynamic Detection of Vulnerable DMA Race Conditions

  • Nov 19, 2025
  • Brian Johannesmeyer +3
  • Conference Article
  • Citations3

A high speed measuring system of yarn tension based on direct memory access

  • Jan 01, 2011
  • Gang Cao +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.