- Book Chapter
- 10.1016/s0927-5452(06)80003-1
Heterogeneous Network-Based Concurrent Computing Systems
- Jan 01, 1995
- Advances in Parallel Computing
- Jack Dongarra
Heterogeneous Network-Based Concurrent Computing Systems
Recently users from both high-performance scientific community and general-purpose applications have shown keen interest in parallel processing due to its higher performance, lower cost, and sustained productivity [148]. To solve a computationally intensive problem efficiently on a cluster of existing computers, distributed computing involves a significantly lower cost factor [156]. Although it is difficult for a distributed computing user to achieve the computational capacity of large massively parallel processors (MPP), it is possible to solve large-size problems by combining a variety of distributed computing resources, connected by high-speed networks. This approach has advantages in terms of flexibility, scalability, and low cost. The advantage of using a cluster of workstations as the computational platform is that a cluster of a large number of workstations is easily available. A disadvantage is that there may be many users running unrelated tasks on the workstations so that the available computing resource for each task fluctuates in an unpredictable manner. Furthermore, communication between workstations is relatively slow. Although the performance of generalized cross-validation (GCV)–based threshold selection scheme is excellent, it is costly from CPU time viewpoint when implemented sequentially. In contrast to the traditional parallel approaches, which rely on specialized parallel machines, present work explores the potential of distributed systems for parallelism. The master/slave model is adopted for control of machines. This chapter is organized as follows. Section 4.1 presents background material and review of work in related area. Section 4.2 presents parallel algorithm of DWT. Basics of PVM programming has been discussed in Sect. 4.3. KeywordsPVMParallel algorithmsSpeedup
Heterogeneous Network-Based Concurrent Computing Systems
Heterogeneous Network-Based Concurrent Computing Systems
Performance driven programmimg models
Most projections for high-performance, massively parallel processors (MPPs) include deep and complex memory hierarchies. Making efficient use of these systems will require making efficient use of these memory hierarchies, without sacrificing the advancements that have been made in algorithms. Efficient programming models were developed for vector computers, particularly the memory system structure, providing high performance. Where are the programming models for MPPs? Much effort has gone into automatic programming systems, such as parallelizing compilers for existing languages and new languages expressing concurrency. Unfortunately, these have rarely led to programs that can achieve near-peak performance. In this paper, we review the issues and some current approaches and suggest some new memory-oriented programming models. The development of these models is essential, because, just as with vector computing, the programming model can strongly influence the new algorithms that are needed for high-performance applications on massively parallel processors.
Read moreTechniques for developing and measuring high performance Web servers over high speed networks
High-performance Web servers are essential to meet the growing demands of the Internet and large-scale intranets. Satisfying these demands requires a thorough understanding of key factors affecting Web server performance. This paper presents empirical analysis illustrating how dynamic and static adaptivity can enhance Web server performance. Two research contributions support this conclusion. First, the paper presents results from a comprehensive empirical study of Web servers (such as Apache, Netscape Enterprise, PHTTPD, Zeus, and JAWS) over high-speed ATM networks. This study illustrates their relative performance and precisely pinpoints the server design choices that cause performance bottlenecks. Once network and disk I/O overheads are reduced to negligible constant factors, the main determinants of Web server performance are its protocol processing path and concurrency strategy. Moreover no single strategy performs optimally for all load conditions and traffic types. Second, we describe the design techniques and optimizations used to develop JAWS, our high-performance adaptive Web server. JAWS is an object-oriented Web server that was explicitly designed to alleviate the performance bottlenecks we identified in existing Web servers. The performance optimizations used in JAWS include adaptive prespawned threading, fixed headers, cached date processing, and file caching. In addition, JAWS uses a novel software architecture that substantially improves its portability and flexibility, relative to other Web servers.
Read moreModified MnO2-Based Cathode for Zinc-Ion Batteries Using Facile Processing and Easily Available Materials
With the fast development of diverse electronics, the demand for safe energy storage systems with high energy density and high stability has increased rapidly. So far, lithium-ion batteries (LIBs) have dominated the market, from small smart electronics to electric vehicles. Nevertheless, LIBs have several limitations, such as high cost, limited raw material resources, high flammability, and harsh environmental impact. Therefore, novel rechargeable batteries which can mitigate these shortcomings must be explored. Aqueous zinc-ion batteries (ZIBs) have shown a promising future. Zinc with high theoretical volumetric capacity (5855 mAh cm-3) and high stability in aqueous electrolytes should enable high-performing, safe, and low-cost batteries. Additionally, the simple assembly of ZIBs in the ambient environment results in low capital costs. Despite these attractive merits, developing cathode materials with high capacity, rate performance, and stability in order to build commercial high-performing ZIBs, remains a great challenge.Cathode materials directly influence the electrochemical performance of ZIBs. Manganese oxide (MnO2) has been investigated as a promising cathode in ZIBs research due to its good specific capacity, low cost and high safety. However, the low ionic conductivity, slow diffusion kinetics, and low stability of MnO2 severely deteriorate the electrochemical performance and cycling stability of zinc-ion batteries. Recently, a series of studies have revealed the high effectiveness of coating on improving MnO2 cathode stability, including polymer coatings, carbon-based material coatings, artificial cathode-electrolyte interfaces, and metal oxide coatings.[1] Nevertheless, the improvement of other parameters, such as the rate performance, could be further improved for practical usage. In this research, we develop a novel modification method combining coating and doping strategies to induce MnO2 cathode with high capacity, high rate performance and low capacity decay after long-term operation. First, we utilize a facile one-step hydrothermal synthesis for Ag-doped α-MnO2. Different amounts of AgNO3 are added into reactors with KMnO4 and MnSO4 (n(Ag):n(Mn)=1:100-4:100) to investigate the influence of doping on capacity and rate performance. Subsequently, the doped tunnel-type MnO2 cathodes are mixed with precursor solutions of Al2O3 and then thoroughly stirred. Finally, an annealing process is conducted to achieve a nanoscale Al2O3 coating on the tunnel type MnO2. The coating configurations influence the electrochemical performance of cathodes;[2] thus, we test precursor solutions with different concentrations (1-3 wt% Al2O3) and annealing temperatures (500-700 °C) to investigate the influence of coating thickness and annealing treatment.With the assistance of scanning and transmission electron microscopes, the nanoscale Al2O3 layer coating on the doped active material is verified. The stable Al2O3 coating protects active materials from dissolution and structural collapse, improving cycling stability and resulting in high capacity retention after long-term operation. Moreover, a significant improvement in rate performance compared to unmodified MnO2 is verified through electrochemical tests. The X-ray photoelectron spectroscopy (XPS) and EPR spectra showed that the doped Ag+ ions lead to oxygen vacancies due to the formation of Ag-O-Mn.[3] The abundant vacancies act as active sites, significantly improving conductivity and ion insertion rates. The ex-situ X-ray diffractometer (XRD) and electrochemical tests are used to investigate the ion insertion mechanism during charging/discharging operations. The Ag-MnO2 @ Al2O3 cathode is synthesized through a facile methodology with environment-friendly materials. The modified cathode exhibits promising performance, especially regarding cycling stability and rate performance, which will be a crucial step for developing high-performing cathodes in aqueous zinc-ion batteries.Reference:[1] Shi, W., Lee, W. S. V., & Xue, J. (2021). Recent development of Mn‐based oxides as zinc‐ion battery cathode. ChemSusChem, 14(7), 1634-1658.[2] Zhou, A., Xu, J., Dai, X., Yang, B., Lu, Y., Wang, L., ... & Li, J. (2016). Improved high-voltage and high-temperature electrochemical performances of LiCoO2 cathode by electrode sputter-coating with Li3PO4. Journal of Power Sources, 322, 10-16.[3] Pu, X., Li, X., Wang, L., Maleki Kheimeh Sari, H., Li, J., Xi, Y., ... & Wu, Y. (2022). Enriching Oxygen Vacancy Defects via Ag–O–Mn Bonds for Enhanced Diffusion Kinetics of δ-MnO2 in Zinc-Ion Batteries. ACS Applied Materials & Interfaces, 14(18), 21159-21172. Figure 1
Read moreA Semi-Lagrangian NWP Model for Real-time and Research Applications: Evaluation in Single- and Multi-Processor Environments
A semi-Lagrangian numerical weather prediction (NWP) model developed for both real-time prediction and for research simulations has been evaluated. The model is second-order in time, high-order (≥ 3) in space, and employs a non-staggered grid in both the horizontal and vertical. A version of the model which is third-order in space was compared with two Eulerian models: the Australian Bureau of Meteorology's current operational regional model which is quasi-second-order in rime and space, and also a new version of this model with a third-order upwind scheme that was developed for use as the Australian Bureau of Meteorology's next operational limited-area NWP model. In a three month trial of twice-daily 48-hour forecasts it was found that both the third-order models were significantly more skillful than the current operational model, as measured by the standard performance statistics such as the S1 skill score (Teweles and Wobus, 1954), and root-mean-square (RMS) errors. The specific implications of this greater accuracy were examined in case studies of severe weather events. The semi-Lagrangian model also bas been adapted to global form and run on a daily basis out to 5 days using archived operational data over a period of almost 6 months. Finally, the semi-Lagrangian model code was parallelized on a workstation cluster and also on a scalable parallel computer, and it was found that the model was well-suited to parallelization on both computer platforms.
Read moreHigh Performance Visualization
Visualization and analysis tools, techniques, and algorithms have undergone a rapid evolution in recent decades to accommodate explosive growth in data size and complexity and to exploit emerging multi- and many-core computational platforms. High Performance Visualization: Enabling Extreme-Scale Scientific Insight focuses on the subset of scientific visualization concerned with algorithm design, implementation, and optimization for use on todays largest computational platforms. The book collects some of the most seminal work in the field, including algorithms and implementations running at the highest levels of concurrency and used by scientific researchers worldwide. After introducing the fundamental concepts of parallel visualization, the book explores approaches to accelerate visualization and analysis operations on high performance computing platforms. Looking to the future and anticipating changes to computational platforms in the transition from the petascale to exascale regime, it presents the main research challenges and describes several contemporary, high performance visualization implementations. Reflecting major concepts in high performance visualization, this book unifies a large and diverse body of computer science research, development, and practical applications. It describes the state of the art at the intersection of scientific visualization, large data, and high performance computing trends, giving readers the foundation to apply the concepts and carry out future research in this area.
Read moreCarrier Dynamics Engineering for High-Performance Electron-Transport-Layer-free Perovskite Photovoltaics
Carrier Dynamics Engineering for High-Performance Electron-Transport-Layer-free Perovskite Photovoltaics
Implementation and analysis of Jacobi iteration based on hybrid programming
With the development of high-speed networks and the multi-core processor technology, the cluster of workstation based on high-speed networks and multi-core processors is becoming the main platform for parallel computing. Jacobi iterative method for solving linear equations is a common method, there are widely range of applications in many areas of science and engineering. This paper parallelizes Jacobi iterative method in process-level using MPI at first, and identifies the most time-consuming part of the program, then parallelizes in thread-level using OpenMP based on shared memory, so that the program can take full advantage of multi-core workstations to reduce the computation time.
Read moreFinal Report for the 10 to 100 Gigabit/Second Networking Laboratory Directed Research and Development Project
The next major performance plateau for high-speed, long-haul networks is at 10 Gbps. Data visualization, high performance network storage, and Massively Parallel Processing (MPP) demand these (and higher) communication rates. MPP-to-MPP distributed processing applications and MPP-to-Network File Store applications already require single conversation communication rates in the range of 10 to 100 Gbps. MPP-to-Visualization Station applications can already utilize communication rates in the 1 to 10 Gbps range. This LDRD project examined some of the building blocks necessary for developing a 10 to 100 Gbps computer network architecture. These included technology areas such as, OS Bypass, Dense Wavelength Division Multiplexing (DWDM), IP switching and routing, Optical Amplifiers, Inverse Multiplexing of ATM, data encryption, and data compression; standards bodies activities in the ATM Forum and the Optical Internetworking Forum (OIF); and proof-of-principle laboratory prototypes. This work has not only advanced the body of knowledge in the aforementioned areas, but has generally facilitated the rapid maturation of high-speed networking and communication technology by: (1) participating in the development of pertinent standards, and (2) by promoting informal (and formal) collaboration with industrial developers of high speed communication equipment.
Read morePlanar junctionless phototransistor: A potential high-performance and low-cost device for optical-communications
Planar junctionless phototransistor: A potential high-performance and low-cost device for optical-communications
Parallel workstation clusters and MPI for sparse systems in computational science
Parallel workstation clusters and MPI for sparse systems in computational science
Design of efficient stamped mirror facets using topography optimisation
Significant cost reduction is required to improve the competitiveness of concentrating solar power. Heliostats make up a significant proportion of the capital cost of solar tower plants, with low-cost, high-performance designs required to meet levelised cost of energy targets. Lightweight stamped mirror facets are seen in leading commercial solar tower plants, with low manufacturing costs and high optical performance. Here, topography optimisation is investigated as a tool for designing lighter and structurally efficient stiffening bead patterns for stamped mirror facets. A case study is presented demonstrating topography optimisation as a promising tool for facet design, with tailored optimisation approaches incorporating wind and gravity load cases, manufacturing constraints, and combined optical-structural objectives. A range of concepts were generated, and their performance assessed to develop a design framework for high-performance supports. Wind loads in the stow position were found to influence and limit weight reduction more than gravity loads during operation. For the various concepts generated by optimisation, a common geometric feature was clearly-defined radial beads from the mounting points to a peripheral bead, driven by the need to overcome high bending stresses. Prototyping stamped structures is typically expensive. A low-cost and accessible rapid-prototyping method for stamped facets was developed using incremental sheet forming. A framework is presented for fabricating facets incorporating optimised supports, with photogrammetry and deflection tests enabling rapid assessment and structural modelling validation.
Read moreDesigning a Temperature Model to Understand the Thermal Challenges of Portable Computing Platforms
Modern high-performance processors are embedded in portable electronics, such as smartphones, self-driving automobiles, and augmented reality wearable. These computing platforms provide real-time direction navigation, high definition video entertainment, real-time sensing and control. One of the major performance limiting factors in these platforms is the poorly designed thermal solution to prevent overheating at the processor transistor junction and at the platform surface. The former limits the maximum operating temperature of transistors to guarantee reliability and lifetime whereas the latter limits the maximum platform surface temperature to ensure user ergonomic comfort. To design effective cooling solutions for high-performance, multi-layer, multi-processor platforms, an accurate and detailed platform level temperature model is needed. In this work, we present a detailed finite volume model to predict the temperature behavior of a tightly packaged, high-performance portable platform, from the system level down to the processor architecture level. We first characterize the thermal response of real-world, representative mobile workloads and perform parametric studies to predict the processor junction and platform surface temperature based on the properties of the platform. We observe that thermal hot spots are sensitive to workload computation characteristics. In particular, for mobile workloads, such as web browsing, computations often occur in short time burst, leading to instantaneous power spikes. This results in temperature spikes as thermal hot spots. To mitigate the thermal issue, modern systems implement dynamic voltage and frequency scaling during high performance and high power scenario to prevent cores from overheating. We validate the proposed temperature model used to predict the maximum temperature at the processor chip and at the platform surface, with measurements from the embedded temperature sensors and an infrared camera. The temperature prediction accuracy for hot spots on the processor chip and on the platform surface is demonstrated to be higher than 90%. With the tool, we show that the physical construction (such as air gaps and material properties) and the floor plan of the multi-layer platform have a profound effect on the system temperature behavior. As the air gap between the different layers of the computing platform increases, the processor temperature rises significantly. Also there exists an optimal thickness setting of the air gap for platform surface thermal spreading. We hope the detailed characterization results, the developed temperature prediction tool, and the insights presented in the paper can motivate advanced thermal management solution for high-performance embedded platforms in the years to come.
Read moreLOW COST AND HIGH PERFORMANCE 5-BIT PROGRAMMABLE PHASED ARRAY ANTENNA AT KU-BAND
We present a low-cost and high-performance 5-bit programmable phased array antenna at Ku-band, which consists of 1-bit reconfigurable radiation structures, digital phase shifters, and a coplanar waveguide feeding network. The 1-bit reconfigurable radiation structure utilizes symmetric geometries and PIN diodes to form stable 180 phase difference. The digital phase shifter provides 168.75 phase difference and together with the radiation structure form a 348.75 phase coverage. The antenna operates between 14.4 and 15.4 GHz, and the overall array contains 24 2 elements with each of them being individually addressable. By changing the states of the diodes and thus adjusting the phase coding sequences of the array, the antenna achieves 0 -60 precise beam scanning at 14.8 GHz, with the sidelobe level, cross-polarization, and gain fluctuation being less than -16 dB, -26 dB, and 2.4 dB, respectively. A prototype was fabricated to verify the design, and the measurement results agree well with simulations. Compared with traditional phased arrays composed of numerous phase shifters and T/R components, the proposed antenna features high performance, high flexibility, low profile, and low cost. The antenna provides a new and feasible solution of wavefront steering and will benefit the various application scenarios.
Read moreViArray standard platforms: Rad-hard structured ASICs for digital and mixed-signal applications
Sandia's radiation-hardened ViArray standard platforms use via-configurable circuits to create quick-turn, low-cost structured ASICs (Application Specific Integrated Circuits). Via-configurable technology enables performance similar to standard-cell ASICs, but with low-volume costs approaching that of programmable logic devices. Due to their significantly lower development cost and shorter production times, ViArray ASICs are rapidly replacing custom ASICs in Sandia's high-reliability digital and analog applications. This paper describes the Eiger ViArray platform, which is optimized for general-purpose digital applications, and the Whistler ViArray platform, which is optimized for mixed-signal instrumentation and state-of-health applications.
Read more