- Research Article
63
- 10.1016/j.comnet.2006.05.003
Modeling and generating realistic streaming media server workloads
- Jun 14, 2006
- Computer Networks
- Wenting Tang + 3 more +3
Modeling and generating realistic streaming media server workloads
Characterization and synthetic generation of streaming access workloads are fundamentally important to the evaluation of Internet streaming delivery systems. GISMO is a toolkit for the generation of synthetic streaming media objects and workloads that capture a number of characteristics empirically verified by recent measurement studies. These characteristics include object popularity, temporal correlation of request, seasonal access patterns, user session durations, user interactivity times, and variable bit rate (VBR) self-similarity and marginal distributions. The embodiment of these characteristics in Gismo enables the generation of realistic and scalable request streams for use in benchmarking and comparative evaluation of Internet streaming media delivery techniques. To demonstrate the usefulness of Gismo, we present a case study that shows the importance of various workload characteristics in evaluating the bandwidth requirements of proxy caching and server patching techniques.
Modeling and generating realistic streaming media server workloads
Modeling and generating realistic streaming media server workloads
Scalable and continuous media streaming on peer-to-peer networks
With the growth of computing power and the proliferation of broadband access to the Internet, media streaming has widely diffused. Although the proxy caching technique is one method to accomplish effective media streaming, it cannot adapt to the variations of user locations and diverse user demands. By using the P2P communication architecture, media streaming can be expected to smoothly react to network conditions and changes in user demands for media-streams. We propose efficient methods to achieve continuous and scalable media streaming system. In our mechanisms, a media stream is divided into blocks for efficient use of network bandwidth and storage space. We propose two scalable search methods and two algorithms to determine an optimum provider peer from search results. Through several simulation experiments, we show that the FLS method can perform continuous media play-out while reducing the amount of search traffic to 1/6 compared with full flooding.
Read moreGeneration of synthetic workloads for multiplayer online gaming benchmarks
We present an approach to the generation of realistic synthetic workloads for use in benchmarking of (massively) multiplayer online gaming infrastructures. Existing techniques are either too simple to be realistic or are too specific to a particular network structure to be used for comparing different networks with each other. Desirable properties of a workload are reproducibility, realism and scalability to any number of players. We achieve this by simulating a gaming session with AI players that are based on behavior trees. The requirements for the AI as well as its parameters are derived from a real gaming session with 16 players. We implemented the evaluation platform including the prototype game Planet PI4. A novel metric is used to measure the similarity between real and synthetic traces with respect to neighborhood characteristics. In our experiments, we compare real trace files, workload generated by two mobility models and two versions of our AI player. We found that our AI players recreate the real workload characteristics more accurately than the mobility models.
Read moreVBR Traffic Shaping for Streaming of Multimedia Transmission
In multimedia applications, media data such as audio and video are transmitted from server to clients via network according to some transmission schedules. Different from the conventional data streams, end-to-end quality-of-service (QoS) is necessary for media transmission to provide jitter-free playback. Therefore, each data packet has been assigned with related timing constraints for transmission. As network resources are allocated exclusively in fixed-size chunks to serve different data streams, it is simple to support constant-bit-rate (CBR) transmission. Grossglauser and Keshav (1996) have investigated the performance of CBR traffic in a large-scale network with many connections and switches. They concluded that the network queuing delay for CBR transmission is less than one cell time per switch even under heavy loading. Besides, resource allocation and admission control are simple as there are no variations in resource requirements. However, media streams are notably variable-bit-rate (VBR) in nature due to the coding and compression technologies applied (Garrett and Willinger, 1994). The average data rate of an MPEG-1 movie is usually less than 25% of its peak data rate. It is inherently at odds with the goals of designing efficient real-time network transmission and admission control mechanisms capable of achieving high resource utilization (Sen et al., 1997). The conventional CBR service model that allocates the peak data rate to transmit the VBR stream would be a waste of bandwidth. Furthermore, it requires a large size of client buffer. To ameliorate this problem, we need a good traffic shaping algorithm to transmit VBR video in a less bursty (i.e., smoother) manner by exploiting different performance measurements. In a multimedia system, we usually measure the performance of a transmission schedule by the following four indices: peak bandwidth, network utilization, initial delay and client buffer.
Read moreA synthetic bursty workload generation method for web 2.0 benchmark
As one of most important characteristics of Web-based systems'workloads, burstiness is gaining more and more attentions. And synthetically generating bursty workloads is a key technique for performance analysis. In this paper, a configurable and intelligible synthetic bursty workload generation method for Web 2.0 benchmark Olio has been proposed based on 2-state Markovian arrival processes (MAP2). By comparing the actual value of index of dispersion for counts (IDC) estimated from system logs with the target value deduced from MAP2 model, we show that our method is more accurate than related work.
Read moreWorkload models and performance evaluation of cloud storage services
Workload models and performance evaluation of cloud storage services
Oracle Workload Intelligence
Analyzing and understanding the characteristics of the incoming workload is crucial in unraveling trends and tuning the performance of a database system. In this work, we present Oracle Workload Intelligence (WI), a tool for workload modeling and mining, as our attempt to infer the processes that generate a given workload. WI consists of two main functionalities. First, WI derives a model that captures the main characteristics of the workload without overfitting, which makes it likely to generalize well to unseen instances of the workload. Such a model provides insights into the most frequent code paths in the application that drives the workload, and also enables optimizations inside the database system that target sequences of query statements. Second, WI can compare the models of different snapshots of the workload to detect whether the workload has changed. Such changes might indicate new trends, regressions, problems, or even security issues. We demonstrate the effectiveness of WI with an experimental study on synthetic workloads and customer-provided application benchmarks.
Read moreMetadata Traces and Workload Models for Evaluating Big Storage Systems
Efficient namespace metadata management is increasingly important as next-generation file systems are designed for peta and exascales. New schemes have been proposed, however, their evaluation has been insufficient due to a lack of appropriate namespace metadata traces. Specifically, no Big Data storage system metadata trace is publicly available and existing ones are a poor replacement. We studied publicly available traces and one Big Data trace from Yahoo! and note some of the differences and their implications to metadata management studies. We discuss the insufficiency of existing evaluation approaches and present a first step towards a statistical metadata workload model that can capture the relevant characteristics of a workload and is suitable for synthetic workload generation. We describe Mimesis, a synthetic workload generator, and evaluate its usefulness through a case study in a least recently used metadata cache for the Hadoop Distributed File System. Simulation results show that the traces generated by Mimesis mimic the original workload and can be used in place of the real trace providing accurate results.
Read moreOn synthetic workloads for multiplayer online games: a methodology for generating representative shooter game workloads
We present approaches to the generation of synthetic workloads for benchmarking multiplayer online gaming infrastructures. Existing techniques, such as mobility or traffic models, are often either too simple to be representative for this purpose or too specific for a particular network structure. Desirable properties of a workload are reproducibility, representativeness, and scalability to any number of players. We analyze different mobility models and AI-based workload generators. Real gaming sessions with human players using the prototype game Planet PI4 serve as a reference workload. Novel metrics are used to measure the similarity between real and synthetic traces with respect to neighborhood characteristics. We found that, although more complicated to handle, AI players reproduce real workload characteristics more accurately than mobility models.
Read moreParallelization of the Array Method Using OpenMP
Shared memory programming and distributed memory programming, are the most prominent ways of parallelize applications requiring high processing times and large amounts of storage in High Performance Computing (HPC) systems; parallel applications can be represented by Parallel Task Graphs (PTG) using Directed Acyclic Graphs (DAGs). The scheduling of PTGs in HPCS is considered a NP-Complete combinatorial problem that requires large amounts of storage and long processing times. Heuristic methods and sequential programming languages have been proposed to address this problem. In the open access paper: Scheduling in Heterogeneous Distributed Computing Systems Based on Internal Structure of Parallel Tasks Graphs with Meta-Heuristics, the Array Method is presented, this method optimizes the use of Processing Elements (PE) in a HPCS and improves response times in scheduling and mapping resource with the use of the Univariate Marginal Distribution Algorithm (UMDA); Array Method uses the internal characteristics of PTGs to make task scheduling; this method was programmed in the C language in sequential form, analyzed and tested with the use of algorithms for the generation of synthetic workloads and DAGs of real applications. Considering the great benefits of parallel software, this research work presents the Array Method using parallel programming with OpenMP. The results of the experiments show an acceleration in the response times of parallel programming compared to sequential programming when evaluating three metrics: waiting time, makespan and quality of assignments.
Read moreGRENCHMARK: A Framework for Analyzing, Testing, and Comparing Grids
Grid computing is becoming the natural way to aggregate and share large sets of heterogeneous resources. With the infrastructure becoming ready for the challenge, current grid development and acceptance hinge on proving that grids reliably support real applications, and on creating adequate benchmarks to quantify this support. However, grid applications are just beginning to emerge, and traditional benchmarks have yet to prove representative in grid environments. To address this chicken-and-egg problem, we propose a middle-way approach: create and run synthetic grid workloads comprising applications representative for today’s grids. For this purpose, we have designed and implemented GRENCHMARK, a framework for synthetic workload generation and submission. The framework greatly facilitates synthetic workload modeling, comes with over 35 synthetic and real applications, and is extensible and flexible. We show how the framework can be used for grid system analysis, functionality testing in grid environments, and for comparing different grid settings, and present the results obtained with GRENCHMARK in our multi-cluster grid, the DAS
Read moreMediSyn
Currently, Internet hosting centers and content distribution networks leverage statistical multiplexing to meet the performance requirements of a number of competing hosted network services. Developing efficient resource allocation mechanisms for such services requires an understanding of both the short-term and long-term behavior of client access patterns to these competing services. At the same time, streaming media services are becoming increasingly popular, presenting new challenges for designers of shared hosting services. These new challenges result from fundamentally new characteristics of streaming media relative to traditional web objects, principally different client access patterns and significantly larger computational and bandwidth overhead associated with a streaming request. To understand the characteristics of these new workloads we use two long-term traces of streaming media services to develop MediSyn, a publicly available streaming media workload generator. In summary, this paper makes the following contributions: i) we model the long-term behavior of network services capturing the process of file introduction and changing file popularity, ii) we present a novel generalized Zipf-like distribution that captures recently-observed popularity of both web objects and streaming media not captured by existing Zipf-like distributions, and iii) we capture a number of characteristics unique to streaming media services, including file duration, encoding bit rate, session duration and non-stationary popularity of media accesses.
Read moreProWGen: a synthetic workload generation tool for simulation evaluation of web proxy caches
ProWGen: a synthetic workload generation tool for simulation evaluation of web proxy caches
Benchmarking a site with realistic workload
The rapidly growing number of Web users and the consequent importance of capacity planning have lead to the development of Web benchmarking tools. One common criticism of this approach, is that synthetic workload produced by Web stressing tools is far from realistic. This paper deals with a benchmarking methodology based on workload characterization generated from log files. A customer behavior model graph (CBMG) was proposed by Mensace, et al., (1999) as workload characterization of an e-commerce site. We discuss how CBMG methodology has a wider field of application and how to use this model to efficiently improve a fully integrated Web stressing tool. We also evaluate the differences between our approach and other models based on different characterizations.
Read moreATLAS grid workload on NDGF resources: analysis, modeling, and workload generation
Evaluating new ideas for job scheduling or data transfer algorithms in large-scale grid systems is known to be notoriously challenging. Existing grid simulators expect to receive a realistic workload as an input. Such input is difficult to provide in absence of an in-depth study of representative grid workloads. In this work, we analyze the ATLAS workload processed on the resources of NDG Facility. ATLAS is one of the biggest grid technology users, with extreme demands for CPU power and bandwidth. The analysis is based on the data sample with ~1.6 million jobs, 1,723 TB of data transfer, and 873 years of processor time. Our additional contributions are (a) scalable workload models that can be used to generate a synthetic workload for a given number of jobs, (b) an open-source workload generator software integrated with existing grid simulators, and (c) suggestions for grid system designers based on the insights of data analysis.
Read more