• Home
  • Search
  • Optimizing sparse tensor times matrix on multi-core and many-core architectures
  • Cite Icon28
  • https://doi.org/10.5555/3018843.3018848Copy DOI Icon

Optimizing sparse tensor times matrix on multi-core and many-core architectures

  • Nov 13, 2016
  • Jiajia Li +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This paper presents the optimized design and implementation of sparse tensor-times-dense matrix multiply (SpTTM) for CPU and GPU platforms. This primitive is a critical bottleneck in data analysis and mining applications based on tensor methods, such as the Tucker decomposition. We first design and implement sequential SpTTM to avoid explicit data transformations between a tensor and a matrix, which is the conventional approach. We further optimize SpTTM on multicore CPU and GPU systems by parallelizing, avoiding locks, and exploiting data locality. Our sequential SpTTM is up to 3.5× faster than the SpTTM from Tensor Toolbox and 1.5× over that from Cyclops Tensor Framework. Our parallel algorithms show 4.1× speedup on multicore Intel Core i7 and 18.8× speedup on NVIDIA K40c GPU over our sequential SpTTM respectively.

Similar Papers
  • Conference Article
  • Citations1

Design and Optimization of High Efficiency Parallel Video Coding System

  • Jun 01, 2012
  • Dongmei Li +2
  • Conference Article
  • Citations24

Pydron: semi-automatic parallelization for multi-core and the cloud

  • Oct 06, 2014
  • Stefan Müller +3
  • Book Chapter
  • Citations110

Hierarchical Place Trees: A Portable Abstraction for Task Parallelism and Data Movement

  • Jan 01, 2010
  • Yonghong Yan +3
  • Research Article
  • Citations7

Parallelization of a Monte Carlo particle transport simulation code

  • Jan 28, 2010
  • Computer Physics Communications
  • P Hadjidoukas +2
  • Research Article

Анализ силовых алгоритмов и инструментов визуализации графа и разработка инструмента визуализации графа

  • Sep 17, 2020
  • S Aubakirov +1
  • Conference Article
  • Citations4

Parallel software for inductance extraction

  • Aug 15, 2004
  • Hemant Mahawar +1
  • Research Article
  • Citations12

FALCON-X: Zero-copy MPI derived datatype processing on modern CPU and GPU architectures

  • May 28, 2020
  • Journal of Parallel and Distributed Computing
  • Jahanzeb Maqbool Hashmi +5
  • Research Article
  • Citations12

Binned k-d Tree Construction for Sparse Volume Data on Multi-Core and GPU Systems.

  • Jan 28, 2021
  • IEEE Transactions on Visualization and Computer Graphics
  • Stefan Zellmann +2
  • Conference Article
  • Citations6

PIMCH: cooperative memory prefetching in processing-in-memory architecture

  • Jan 22, 2018
  • Sheng Xu +3
  • Conference Article
  • Citations69

Balancing processor loads and exploiting data locality in N-body simulations

  • Jan 01, 1995
  • Ioana Banicescu +1
  • Conference Article
  • Citations5

Accelerated AC contingency calculation on commodity multi-core SIMD CPUs

  • Jul 01, 2014
  • Tao Cui +3
  • Research Article

A memory and blocks of threads management in parallel computation using CUDA architecture

  • Dec 01, 2010
  • Jacek Widuch
  • Research Article
  • Citations1

Boosting Python performance on Intel processors: a case study of optimizing music recognition

  • Nov 13, 2016
  • Yuanzhe Li +1
  • Dissertation
  • Citations1

Applications on emerging paradigms in parallel computing

  • Apr 30, 2012
  • Abhinav Sarje
  • Conference Article
  • Citations19

First Impressions of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchip for Scientific Workloads

  • Jan 11, 2024
  • Nikolay A Simakov +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.