- Research Article
1
- 10.1016/0020-0190(90)90049-4
An algorithm for load balancing in multiprocessor systems
- Aug 01, 1990
- Information Processing Letters
- Michael C Loui + 1 more +1
An algorithm for load balancing in multiprocessor systems
Nginx is a commonly used and free open-source web server that is used as a reverse proxy server, load balancer and HTTP cache. It consumes less memory and can handle more clients with less number of processes. Nginx provides users with five predefined load balancing algorithms. However, most of these algorithms are static and some of the load balancing rules are inefficient. In order to make the load of a cluster more stable under high concurrent requests, we developed a Dynamic Load Balancing (DLB) algorithm that uses Nginx as a network security control panel to provide load balancing for a cluster of backend servers. The DLB algorithm is based on the weighted round robin module of Nginx, Logistic Regression and Maximum Likelihood Estimation (MLE) algorithm. It handles the situation of high concurrent requests and reduces the probability of omitted or under-reported incident and status. We also propose a Hybrid Load Balance Method (HLBM) that incorporates the DLB algorithm and evolutionary computing to further improve the performance. We have conducted limited experiment by using dynamic load balancing algorithm. We will complete to develop the Hybrid Load Balance method and conduct experiments for HLBM.
An algorithm for load balancing in multiprocessor systems
An algorithm for load balancing in multiprocessor systems
CFDA parallel Method for Adaptive Unstructured Grids with Optimum Static Grid Repartitioning
CFDA parallel Method for Adaptive Unstructured Grids with Optimum Static Grid Repartitioning
Load balancing algorithms for an extended hypercube
Reduction of the execution time of a job through equitable distribution of work load among the processors in a distributed system is the goal of load balancing. Performance of static and dynamic load balancing algorithms for the extended hypercube, is discussed. Threshold algorithms are very well-known algorithms for dynamic load balancing in distributed systems. An extension of the threshold algorithm, called the multilevel threshold algorithm, has been proposed. The hierarchical interconnection network of the extended hypercube is suitable for implementing the proposed algorithm. The new algorithm has been implemented on a transputer-based system and the performance of the algorithm for an extended hypercube is compared with those for mesh and binary hypercube networks.
Read moreA decentralized algorithm for dynamic load balancing with file transfer
A decentralized algorithm for dynamic load balancing with file transfer
Parallel and Distributed Computing Techniques, Selection of papers from ISPDC 2008
Dear SCPE Reader, We present a selection of papers which are extensions of papers presented at the 7-th International Symposium on Parallel and Distributed Computing, 1–5 July 2008, in Krakow, Poland. The motivation for publishing the selection in the SCPE Journal was, on the one hand, to present the flavour of the research reported at the conference and on the other hand to present some of the most relevant topics currently focused on the research on parallel and distributed computing in general. The selection contains only 6 papers out of about 60 presented at the conference, and thus, is far from covering all relevant topics represented at the ISPDC 2008. This is because not all of the invited authors were patient enough to accept a fairly long paper publishing process. Nevertheless, we hope that the presented papers will bring you closer to the research covered by the ISPDC conferences and will encourage you to participate in future ISPDC editions. The first paper ``The Impact of Workload Variability on Load Balancing Algorithms'' is by Marta Beltran and Antonio Guzman from King Juan Carlos University in Spain. It concerns an important topic of load balancing in cluster systems, namely adaptativity of the load balancing algorithms to changes of the workload in the system. Adequate accounting for additional load in the hosting system is of great relevance for correct optimization effects. The paper presents a thorough formal analysis of the workload variability metrics and their influence on the quality of load balancing algorithms. Four basic activities appearing in load balancing algorithms are identified, and based on them some algorithmic solutions are proposed to correctly deal with workload variability in system load balancing. The problem of dynamic load balancing algorithms robustness has been discussed. Two different robustness metrics sensitive to the applied type of opimization: local task-oriented or a global one enable selecting task remote execution or migration as load balancing operations. The proposed approach is illustrated with experiments. The second paper ``Model-Driven Engineering and Formal Validation of High-Performance Embedded Systems'' is by Abdoulaye Gamatie, Eric Rutten, Huafeng Yu, Pierre Boulet, Jean-Luc Dekeyser, from University of Lille and INRIA in France. The paper is concerned with a very advanced methodology of designing correct parallel embedded systems for intensive data-parallel computing. In their previous research, the authors of the paper designed the GASPARD embedded system design framework. It is based on the hardware/software co-design approach through model-driven engineering. The framework is based on an UML-like model specification language in which hardware and software elements are modelled using a component approach with special mechanisms for repetitive structures. This paper tries to combine the modelling framework of GASPARD with the mechanisms of synchronous languages to achieve design verifiability provided for such languages. The paper shows how GASPARD models can be translated into synchronous models based on data flow equations in order to formally check their correctness. The proposed approach is illustrated with an example of a video processing system. The third paper ``Relations Between Several Parallel Computational Models'' is by Stephan Bruda and Yuanqiao Zhang from Bishop’s University in Canada. The paper is concerned with theoretical aspects of shared memory systems described by the parallel random access machine PRAM model and aims in studying performance properties of different types of PRAM systems. The attention is focussed on analysing the computational power of two more sophisticated PRAM models (Combining CRCW and Broadcast Selective Reduction), which include data reduction in case of concurrent writes. The paper shows that these two models have equivalent computational power, which is a new result comparing the existing literature. The performance of both models applied to reconfigurable multiple bus machines was studied as a possible architectural solution for current VLSI processor implementations. It was shown that in such systems under reasonable assumptions concurrent-write does not enhance performance comparing the exclusive-write model. Another result important for the VLSI technology is that the Combining CRCW PRAM model (in which data of concurrent writes are arthmetically or logically combined before write) and the exclusive-write on directed reconfigurable busses perform in equivalent way under strong real-time requirements. The fourth paper ``Experiences with Mesh-Like Computations Using Prediction Binary Trees'' is by Gennaro Cordasco, Biagio Cosenza, Rosario de Chiara, Ugo Erra and Vittorio Scarano from the University ``degli Studi'' of Salerno and the University ``degli Studi della Biasilicata'' of Potenza in Italy. The paper concerns optimization methods for mesh-like computations in clusters of processors. The computations are perfomed assuming a phase-like program execution control using a tiling approach which reduces inter-processor communication. A temporal coherence is also assumed, which means that task sizes provide similar execution times in consecutive phases. Temporary coherent computations are structured in a Prediction Binary Tree, in which leaves represent computing tiles to be mapped to processors. A phase-by-phase semi-static load balancing is introduced to the scheduling algorithm. The scheduling algorithm is equipped with a predictor, which estimates the computation time of next phase tiles based on previous execution times and modifies the tiles to achieve balanced execution in phases. For this, two heuristics are used to leverage on data locality in processors. The proposed approach is illustrated by the example of interactive rendering with Parallel Ray Tracing algorithm. The fifth paper ``The Influence of the IBM pSeries Servers Virtualization Mechanism on Dynamic Resource Allocation in AIX 5L'' is by Maciej Mlynski from ASpartner Limited in Poland. The paper concerns a very up-to-date problem of system virtualization and presents the results of research carried on IBM pSeries servers. IBM is strongly developing the virtualization technique especially on IBM pSeries servers enabling an improved and flexible sharing of system resources between applications. The paper investigates novel facilities for dynamic resource management such as micro-partitioning and partition load manager. They enable dynamic creation of workload logical partitions of system resources and their dynamic mangement. It includes run-time resource re-alocation between logical partitions including setting of sharing specifications as well as run-time adding/removing/setting parameters of resources in the system. It remains an open question how to properly tune parameters of the operating system using the provided virtualization facilities to obtain the best efficiency for a given application program. The paper presents the results of experiments which study the effects of tuning the disk subsystem parameters under the IBM AIX 5L operating system with the use of the provided virtualization facilities on the resulting application execution performance. The results show that even small deterioration in the resource pool status requires an immediate adaptation of the operating system parameters to maintain the required performance. The sixth paper ``HeteroPBLAS: A Set of Parallel Basic Linear Algebra Subprograms Optimized for Heterogeneous Computational Clusters'' is by Ravi Reddy, Alexey Lastovetsky and Pedro Alonso from University College Dublin in Ireland and Polytechnic University of Valencia in Spain. The paper concerns the methodology for parallelization of linear algebra computations for execution in heterogeneous cluster environments. The design of the HeteroPBLAS library (Parallel Basic Linear Algebra Subprograms) for heterogeneous computational clusters is presented. The main contribution of the paper is the automation of the parallelization and optimization of the PBLAS, which is done by means of a special user interface and the underlying set of functions. An important element is here a performance model that is based on program code instrumentation, which determines parameters of the application and the executive heterogeneous platform relevant for execution performance of parallel code. The parameter values specified for or returned by execution of the performance model functions are next used for generation and optimal mapping of the parallel code of the library subroutines. The proposed approach is illustrated by experimental results of execution of optimized HeteroPBLAS programs on homogeneous and heterogeneous computing clusters. Marek Tudruj
Read moreA physical particle and plane framework for load balancing in multiprocessors
Different models for load balancing have been proposed before, each of which has its own features and advantages when considered for a specific scenario. Yet, nearly all of the existing techniques have assumed an oversimplified model of the system which is often not the case of the real world. In this paper, a new gradient based algorithm for dynamic load balancing on multiprocessors is proposed. This algorithm is an analogy of a classical physical model of a Particle & Plane system which operates based on the classic laws of physics dictated by the nature.
Read moreDesign and implementation of dynamic load balancing algorithms for rollback reduction in optimistic PDES
In an optimistic (time-warp) parallel simulation, local clocks with different logical processes must advance at the same rate in order to reduce the number of rollbacks. In this paper, we propose two algorithms for dynamic load balancing which reduce the number of rollbacks in an optimistic parallel discrete event simulation (PDES) system. The first algorithm is based on the load transfer mechanism between logical processes, while the second algorithm, which is based on the principle of an evolutionary strategy, migrates logical processors between several pairs of physical processors. We have implemented both of these algorithms on a cluster of heterogeneous workstations and studied their performance. The experimental results show that the algorithm based on the load transfer is effective when the grain size is larger than 10 ms, and the algorithm based on the process migration yields good performance for grain sizes of 20 ms or larger. In both of these cases, the average speed-up ranges between 1 and 2 using four processors, when the computation grain-size is within the range 7 to 50 ms. The reduction in rollback messages as a percentage of the total number of messages due to the algorithms is, however, around 4 to 8%.
Read moreA comparative study and analysis of agent based monitoring and fuzzy load balancing in distributed systems
Load balancing algorithms are designed to achieve efficient resource utilisation in large-scale distributed systems. The primary goal of designing load balancing algorithms is to balance the overall workload among all the nodes in distributed systems to improve the performance. This paper presents an in-depth study and detailed comparative analysis of different load monitoring and balancing algorithms employing fuzzy logic and mobile agents. A detailed qualitative as well as quantitative analysis of algorithmic performances is included. The analytical survey presented in this paper would enable understanding of applications of different algorithms under different computing environments. Furthermore, this paper proposes a hybrid architecture combining fuzzy logic and mobile agents for load balancing and monitoring in distributed systems.
Read moreResearch on Online Interactive Teaching Platform of College English Combined with Semantic Association Network Modeling
The emergence of semantic association networks has injected a new impetus for the development of online English teaching and provided a new model reference for the design of online education platforms. In this paper, the research and design of an online interactive teaching platform for college English draws on the algorithmic advantages of the semantic associative network model and utilizes the self-operation of the semantic associative network to realize the functions of autonomous addition, deletion, modification, and checking. The text semantic similarity is predicted by word embedding model, convolutional neural network, and other algorithms so as to better achieve the integration of teaching resources, connecting English knowledge and highlighting the teaching focus in the online teaching process of college English. Dynamic load balancing algorithms are used to solve the problems of short-term surges in the number of visits and the concentration of call requests, and the optimization of load balancing algorithms is further realized through genetic algorithms to finally complete the design of the online teaching interactive platform. Comparison experiments concluded that the semantic association network proposed in this paper could hold a more stable repair effect when cleaning inconsistent data in the dataset, highlighting the effectiveness of the semantic association network model in this paper. The online interactive teaching platform designed in this paper also performs well in the performance test, with only a 0.01% abnormality rate in the concurrency performance test, and the load balancing ability test also achieves the expected effect.
Read moreAdaptive Neuro Fuzzy Interference and PNN Memory Based Grey Wolf Optimization Algorithm for Optimal Load Balancing
In recent years, cloud computing provides a spectacular platform for numerous users with persistent and alternative varying requirements. In the cloud environment, security and service availability are the two most significant factors during the data encryption process. For providing optimal service availability, it is necessary to establish a load balancing technique that is capable of balancing the request from diverse nodes present in the cloud. This paper aims in establishing a dynamic load balancing technique using the APMG approach. Here in this paper, we integrated adaptive neuro-fuzzy interference system-polynomial neural network as well as memory-based grey wolf optimization algorithm for optimal load balancing. The memory-based grey wolf optimization algorithm is employed to enhance the precision of ANFIS-PNN and to maximize the locations of the membership functions respectively. Also, two significant factors namely the turnaround time and CPU utilization involved in optimal load balancing scheme are evaluated. Finally, the performance evaluation of the proposed MG-ANFIS based dynamic load balancing approach is compared with various other load balancing approaches to determine the system performances.
Read moreAn Improved Multimedia Conference System with Load Balance
With the development of economics and the globalization of trade, the Multimedia Conference System is now widely used and the number of its users increases exponentially. But a single server can't handle the high concurrent user requests effectively. In this paper, we improved the Multimedia Conference System with load balance technology. With the load balance technology, the robustness and availability of the system can be enhanced significantly. Keywords-load balance, cluster, Nginx, Multimedia Conference System I. INTRODUCTION With the explosive popularity of the Internet and the World Wide Web, most popular web sites confronted with a problem that the server overloaded caused by the increasing number of concurrent accesses. In order to process the user requests timely, increase the network throughput and improve the quality of service, it's necessary to upgrade the hardware and software of the sever. But the the hardware is expensive and non-scalability. In this case, server cluster system appears. The sever cluster system refers to a server group which is composed of more than one homogeneous or heterogeneous severs and can provides services that are transparent for the external users. And it becomes a key point that how to achieve a reasonable distribution of load between multiple servers and avoid appearing one server with full load while other servers with little load. In this situation, the load balance technology was born. In recent years, cluster technology and load balance technology got fully developed and they have significant effect on solving the problem of overloading. As a web service, the Multimedia Conference System also has to face the problem that the overload of the server and uneven distribution of resources when there are too many users access the system. Using cluster and load balance technology, the requests the system received can be evenly distributed on each server, thus we can reduce the response time of the requests and improve the resource utilization. The Multimedia Conference System proposed in this paper contains Client Layer, Load Balancer Layer, Server Cluster Layer and Media Server Layer. The Client Layer receives the requests of the user including creating a conference and joining an existing conference. The Load Balancer Layer redirect the user's request to one of the application servers in the cluster according to the load balance algorithm and the different parameters from the Client Layer. The application servers in Server Cluster Layer call the Media Server resources and return them to the Client Layer. This implements a basic process of the system. The Load Balancer Layer contains four modules including User Authentication, Decision-Making, Redirecting and Log Parsing. With the Load Balancer Layer and its load balance algorithm, the user's request can be distributed on different servers and those who want to join the same conference can be allocated to the same application server. Thus we can minimize the response time of the request, improve the resource utilization and also ensure the conference can be held correctly. The rest of this paper is organized as follows: we describe related work in section II. Section III presents the architecture of the Multimedia Conference System. In section IV, we display the implementation of the Load Balancer Layer in detail. Some experiments of key features of load balancer in section V. We conclude in section VI.
Read moreA reconfigurable MPSoC-based QAM modulation architecture
QAM is a widely used multi-level modulation technique, with a variety of applications in data radio communication systems. Most existing implementations of QAM-based systems use high levels of modulation in order to meet the high data rate constraint of emerging applications. This work presents the architecture of a highly-parallel MPSoC-based QAM modulator that offers multi-rate modulation. The proposed MPSoC architecture is modular and provides flexibility via dynamic reconfiguration of the QAM, offering high data rates (more than 1 Gbps), even at low modulation levels (16-QAM). Furthermore, the proposed QAM implementation integrates a hardware-based resource allocation algorithm for dynamic load balancing.
Read moreLoad balancing and switch scheduling
Packet switching remains one of the bottlenecks in building fast Internet routers. Load balancing and switch scheduling are two important algorithms in the effort to maximize the throughput and minimize the latency of these packet switches. A load balancing algorithm regulates the traffic to conform to the service rates while a switch scheduling algorithm allocates the service rates adaptive to the arrival patterns. Many existing load balancing and switch scheduling algorithms are very similar. We show that load balancing and switch scheduling systems are dual systems based on the linear queue dynamics approximation. This allows us to cast a load balancing problem as a scheduling problem, and vice versa. We further show an example of designing a new algorithm for load balancing using an existing scheduling algorithm based on the duality. The duality perspective also allows us to solve unknown problems. We find the entropy rate of the randomized bandwidth allocation system with linear queue dynamics based on the knowledge of the entropy rate of the randomized load balancing system. For the general case, we find both an upper and a lower bound on the entropy rate. The joint use of dual load balancing and switch scheduling algorithms leads to performance gains as we show using mean field analysis.
Read moreDesign and Performance Evaluation of Queue-and-Rate-Adjustment Dynamic Load Balancing Policies for Distributed Networks
In this paper, we classify the dynamic distributed load balancing algorithms for heterogenous distributed computer systems into three policies: queue adjustment policy (QAP), rate adjustment policy (RAP), and queue and rate adjustment policy (QRAP). We propose two efficient algorithms, referred to as rate-based load balancing via virtual routing (RLBVR) and queue-based load balancing via virtual routing (QLBVR), which belong to the above RAP and QRAP policies, respectively. We also consider algorithms estimated load information scheduling algorithm (ELISA) and perfect information algorithm, which were introduced in the literature, to implement QAP policy. Our focus is to analyze and understand the behaviors of these algorithms in terms of their load balancing abilities under varying load conditions (light, moderate, or high) and the minimization of the mean response time of jobs. We compare the above classes of algorithms by a number of rigorous simulation experiments to elicit their behaviors under some influencing parameters, such as load on the system and status exchange intervals. We also extend our experimental verification to large scale cluster systems such as a mesh architecture, which is widely used in real-life situations. From these experiments, recommendations are drawn to prescribe the suitability of the algorithms under various situations
Read moreAn Empirical Study of Data Partitioning and Replication in Parallel Simulation
Two of the important design decisions in developing a parallel program are how the data space is to be partitioned and how much data repli- cation there should be among the various processes. In this paper, we propose a strategy for handling data distribution and replication for the problem of proximity detection in a parallel combat simula- tion. The strategy is compared with similar strate- gies proposed for other distributed simulations. Performance results are presented using JPL's Time Warp Operating System with 70 processors on a BBN Butterfly GP1000. It is found that this dis- tributed proximity detection algorithm generalizes well and has acceptable performance. Our vehicle for parallel execution is the Time Warp Operating System (TWOS) under development at the Jet Propulsion Laboratory since 1984 (l). TWOS is based on the theory of virtual time, which relies on optimistic scheduling, process roll- back, and message cancellation as its primary vehi- cles to ensure correct parallel execution (2). JPL's implementation of TWOS is quite advanced, incor- porating algorithms for space management (3) and dynamic load balancing (4), both of which are de- fined entirely in terms of virtual time. Although virtual time can be used for distributed database concurrency control, real time systems, and other applications where there is a natural interpretation of virtual time, JPL's implementation is limited to supporting distributed simulations. This paper will concentrate on performance experi- ments designed to understand some of the differ- ences between parallel and sequential implementa- tions of the same simulation. Understanding the differences between two such implementations is important for several reasons. First, to the extent such differences might be reduced to a mere transla- tion of a sequential model into a parallel one, a compiler could be designed to accomplish the task. Secondly, the differences might shed light on the difficulty of writing an efficient distributed simula- tion, given the biases inherited from years of se- quential programming. Thirdly, the results will guide future efforts at developing distributed simu- lations, efforts which are currently underway at JPL and other laboratories.
Read more