• Home
  • Search
  • Experience with building a commodity Intel-based ccNUMA system
  • Cite Icon13
  • https://doi.org/10.1147/rd.452.0207Copy DOI Icon

Experience with building a commodity Intel-based ccNUMA system

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Commercial cache-coherent nonuniform memory access (ccNUMA) systems often require extensive investments in hardware design and operating system support. A different approach to building these systems is to use Standard High Volume (SHV) hardware and stock software components as building blocks and assemble them with minimal investments in hardware and software. This design approach trades the performance advantages of specialized hardware design for simplicity and implementation speed, and relies on application-level tuning for scalability and performance. We present our experience with this approach in this paper. We built a 16-way ccNUMA Intel system consisting of four commodity four-processor Fujitsu® Teamserver™ SMPs connected by a Synfinity™ cache-coherent switch. The system features a total of sixteen 350-MHz Intel® Xeon™ processors and 4 GB of physical memory, and runs the standard commercial Microsoft Windows NT® operating system. The system can be partitioned statically or dynamically, and uses an innovative, combined hardware/software approach to support application-level performance tuning. On the hardware side, a programmable performance-monitor card measures the frequency of remote-memory accesses, which constitute the predominant source of performance overhead. The monitor does not cause any performance overhead and can be deployed in production mode, providing the possibility for dynamic performance tuning if the application workload changes over time. On the software side, the Resource Set abstraction allows application-level threads to improve performance and scalability by specifying their execution and memory affinity across the ccNUMA system. Results from a performance-evaluation study confirm the success of the combined hardware/software approach for performance tuning in computation-intensive workloads. The results also show that the poor local-memory bandwidth in commodity Intel-based systems, rather than the latency of remote-memory access, is often the main contributor to poor scalability and performance. The contributions of this work can be summarized as follows: • The Resource Set abstraction allows control over resource allocation in a portable manner across ccNUMA architectures; we describe how it was implemented without modifying the operating system. • An innovative hardware design for a programmable performance-monitor card is designed specifically for a ccNUMA environment and allows dynamic, adaptive performance optimizations. • A performance study shows that performance and scalability are often limited by the local-memory bandwidth rather than by the effects of remote-memory access in an Intel-based architecture.

Similar Papers
  • Conference Article

Design and Implementation of BIOS for Godson-3A Interconnections

  • May 01, 2011
  • Yuhui Gao +4
  • Conference Article
  • Citations3

An improved data communication mechanism for a SOC hardware/software co-emulation environment

  • Jul 01, 2009
  • A.W Ruan +4
  • Book Chapter

Technologies for High-Performance Computing in the Next Millennium

  • Jan 01, 2000
  • Dave Turek
  • Conference Article
  • Citations7

A tool for enforcing system structure

  • Jan 01, 1973
  • John R White +1
  • Research Article

A QR-Enabled Multi-Participant Quiz System for Educational Settings with Configurable Timing

  • Oct 22, 2025
  • Applied System Innovation
  • Junjie Li +5
  • Conference Article
  • Citations3

Wireless granary temperature measurement node with function of low voltage detection

  • May 01, 2016
  • Wang Fudong +4
  • Book Chapter
  • Citations3

A Computer Vision Based Machine for Walnuts Sorting Using Robot Operating System

  • Nov 12, 2016
  • Truong Tran +3
  • Conference Article
  • Citations3

Energy consumption optimization of the Total-FETI solver and BLAS routines by changing the CPU frequency

  • Jul 01, 2016
  • David Horak +4
  • Single Book
  • Citations267

Industrial Cloud-Based Cyber-Physical Systems

  • Jan 01, 2014
  • Armando W Colombo +1
  • PDF
  • Research Article
  • Citations10

A Social Environmental Sensor Network Integrated within a Web GIS Platform

  • Nov 21, 2017
  • Journal of Sensor and Actuator Networks
  • Yorghos Voutos +3
  • Conference Article
  • Citations503

State of the art of virtual reality technology

  • Mar 01, 2016
  • Christoph Anthes +3
  • Research Article
  • Citations7

Performance analysis and tuning for a single-chip multiprocessor DSP

  • Jan 01, 1997
  • IEEE Concurrency
  • Jihong Kim +1
  • Book Chapter

Extended Overhead Analysis for OpenMP Performance Tuning

  • Jan 01, 2003
  • Chen Yongjian +2
  • Conference Article

Hardware / software codesign and implementation for secure NFC applications

  • May 01, 2015
  • Subutay Giray Başkır +1
  • Research Article
  • Citations62

Multi-objective performance optimisation for model predictive control by goal attainment

  • Jul 01, 2010
  • International Journal of Control
  • Vasileios Exadaktylos +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.