- Single Book
200
- 10.1016/c2009-0-62200-5
Multidimensional Signal, Image, and Video Processing and Coding
- Jan 01, 2012
- John W Woods
Multidimensional Signal, Image, and Video Processing and Coding
The computational complexity issue is critical for present and future video applications implemented by relatively new video coding standards, such as the H.264/AVC (Advanced Video Coding), which has a large number of coding modes. One of the main reasons for the importance of providing an efficient complexity control in video coding applications is a strong need to decrease the encoding/decoding computational complexity, especially when the encoding and/or decoding devices are resource-limited, such as portable devices. In turn, efficient complexity control enables reducing the video coding processing time and enables saving power resources during the encoding and/or decoding process. Since the recent dramatic progress in the development of multimedia technologies has made portable devices widespread everywhere, especially in order to provide or receive real-time video contents, the need to enhance the computational complexity control in video coding applications is expected to be further significantly increased as a function of the dramatic increase in the mobile/portable device penetration into the every-day life environment. In this chapter, the authors perform a comprehensive review of the recent advances in computational complexity techniques for video coding applications. This chapter will not only summarize the recent advances in this field, but will also provide explicit directions for the design of the future complexity-aware video coding applications.
Multidimensional Signal, Image, and Video Processing and Coding
Multidimensional Signal, Image, and Video Processing and Coding
Spatio-temporal segmentation and object tracking
Visual information is taking a predominant place in our society. With the advent of digital technologies, visual information will find new applications in domains ranging from communication, commerce to entertainment. New functionalities will be required that permit an extended interaction with the visual information. In particular, this is the case for digital television and the related problem of video sequence compression. The extent of possible interactions depends on the manner in which the visual information is represented. Up to now, the canonical representation is used. Also referred to as the waveform representation, it is a technical artifact of the image capture procedure. This representation severely restricts the functionalities available to the end-user. The latter is unable to freely manipulate or customize the received visual information. In this dissertation, we propose to describe the visual information through a semantically meaningful representation. This representation directly derives from the scene content, which is decomposed in terms of its constituting objects. Consequently, the viewing process is totally disconnected from the image capture procedure. This permits full interactivity with the visual information, leading to enhanced functionalities for the end-user. The essence of this dissertation is to automatically define the objects forming the scene arid to automatically track them through the video sequence. Two main issues are identified: the initial segmentation of the objects and their tracking. The first issue deals with segmenting the scene into its constituent objects. This has to be performed on the basis of the information available in two consecutive frames. In order to solve the problem, a split-and-merge approach is used. First, the image is segmented into small, spatio-temporally homogeneous regions. This is achieved through a top-down approach where spatial, temporal and change information is combined. These regions are used as a starting point to define the objects forming the scene. A bottom-up approach is used which iteratively merges the regions. The propensity of the regions to form an object is assessed in terms of both spatial and temporal information. The second issue deals with segmenting and tracking the objects in the successive frames. The coherence between the successive segmentations is ensured by using past and current information. Also, the different objects composing the scene are identified throughout the video sequence. This identification relies on temporal, spatial and spatio-temporal features of the objects. The proposed representation of the visual information finds a natural application in video coding. A second generation video coding scheme is presented which combines compression efficiency with extended functionalities. In particular, scalable coding is achieved in terms of both scene content and object coding quality.
Read moreA highly efficient multiplication-free binary arithmetic coder and its application in video coding
A novel and highly efficient algorithm of multiplication-free binary arithmetic coding is proposed. Our proposed method relies on simple table lookups for performing the computationally critical operations of interval subdivision and probability estimation. Moreover, the underlying design principle provides a great flexibility for serving the different needs of all kind of coding applications where binary or binarized data have to be processed. A binary arithmetic coder of the type described in this paper has become part of the CABAC entropy coding scheme of the emerging H.264/AVC video coding standard. Experiments using this binary arithmetic coder in its native video coding environment demonstrate a superior coding efficiency as well as a significantly higher throughput rate in comparison to the MQ coder, which is currently being considered state-of-the-art in fast binary arithmetic coding.
Read moreReal-time and parallel SHVC hybrid codec AVC to HEVC decoder
Scalable High efficiency Video Coding (SHVC) is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). One of the key novelties introduced by SHVC is that it enables hybrid codec scalability. This basically means that the video layers can be encoded with different video standards providing backward compatibility between codecs. In this paper, we propose a software parallel SHVC decoder in hybrid codec scalability configuration. The proposed design consists of an Advanced Video Coding (AVC) decoder for the Base Layer (BL) and a HEVC decoder for the Enhanced Layer (EL). In order to perform Inter Layer Prediction (ILP), a communication of decoding states and outputs is established between the two decoders. While the native frame based parallelism is still allowed within the two decoders, the proposed design also enables the use of frame based parallelism between the two decoders. The proposed software design enables a real time decoding of the HEVC EL at 2160p60 while the AVC base layer is decoded at 1080p60 for ×2 spatial scalability.
Read moreVLSI architectures design for encoders of High Efficiency Video Coding (HEVC) standard
The growing popularity of high resolution video and the continuously increasing demands for high quality video on mobile devices are producing stronger needs for more efficient video encoder. Concerning these desires, HEVC, a newest video coding standard, has been developed by a joint team formed by ISO/IEO MPEG and ITU/T VCEG. Its design goal is to achieve a 50% compression gain over its predecessor H.264 with an equal or even higher perceptual video quality. Motion Estimation (ME) being as one of the most critical module in video coding contributes almost 50%-70% of computational complexity in the video encoder. This high consumption of the computational resources puts a limit on the performance of encoders, especially for full HD or ultra HD videos, in terms of coding speed, bit-rate and video quality. Thus the major part of this work concentrates on the computational complexity reduction and improvement of timing performance of motion estimation algorithms for HEVC standard. First, a new strategy to calculate the SAD (Sum of Absolute Difference) for motion estimation is designed based on the statistics on property of pixel data of video sequences. This statistics demonstrates the size relationship between the sum of two sets of pixels has a determined connection with the distribution of the size relationship between individual pixels from the two sets. Taking the advantage of this observation, only a small proportion of pixels is necessary to be involved in the SAD calculation. Simulations show that the amount of computations required in the full search algorithm is reduced by about 58% on average and up to 70% in the best case. Secondly, from the scope of parallelization an enhanced TZ search for HEVC is proposed using novel schemes of multiple MVPs (motion vector predictor) and shared MVP. Specifically, resorting to multiple MVPs the initial search process is performed in parallel at multiple search centers, and the ME processing engine for PUs within one CU are parallelized based on the MVP sharing scheme on CU (coding unit) level. Moreover, the SAD module for ME engine is also parallelly implemented for PU size of 32×32. Experiments indicate it achieves an appreciable improvement on the throughput and coding efficiency of the HEVC video encoder. In addition, the other part of this thesis is contributed to the VLSI architecture design for finding the first W maximum/minimum values targeting towards high speed and low hardware cost. The architecture based on the novel bit-wise AND scheme has only half of the area of the best reference solution and its critical path delay is comparable with other implementations. While the FPCG (full parallel comparison grid) architecture, which utilizes the optimized comparator-based structure, achieves 3.6 times faster on average on the speed and even 5.2 times faster at best comparing with the reference architectures. Finally the architecture using the partial sorting strategy reaches a good balance on the timing performance and area, which has a slightly lower or comparable speed with FPCG architecture and a acceptable hardware cost
Read moreA Fast CU Depth Decision Algorithm Based on Moving Object Detection for High Efficiency Video Coding
Compared with Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) standard has higher compression ratio under the same reconstructed video quality. The improvement of coding ability depends on the introduction of many new coding technologies, such as flexible quadtree structure and more prediction modes. However, these coding tools also bring great computational complexity, which brings challenges to the real-time application of HEVC. Therefore, reducing the computational complexity of HEVC is a very meaningful topic. In this paper, firstly, we use Binary Sum of Absolute Difference (BSAD) to divide Coding Tree Unit (CTU) into three different types, namely static region, moving object boundary and moving object interior. Then, we utilize the depth distribution characteristics of different types to predict the depth range of Coding Tree Unit (CTU). Experimental results show that the proposed algorithm can effectively avoid unnecessary depth traversal, save 48.75% of coding time on average, and only increase the bitrate by 0.98%.
Read moreAn adaptive scan of high frequency subbands for dyadic intra frame in MPEG4-AVC/H.264 scalable video coding
This paper develops a new adaptive scanning methodology for intra frame scalable coding framework based on a subband/wavelet(DWTSB) coding approach for MPEG-4 AVC/H.264 scalable video coding (SVC). It attempts to take advantage of the prior knowledge of the frequencies which are present in different higher frequency subbands. We propose dyadic intra frame coding method with adaptive scan (DWTSB-AS) for each subband as traditional zigzag scan is not suitable for high frequency subbands. Thus, by just modification of the scan order of the intra frame scalable coding framework of H.264, we can get better compression. The proposed algorithm has been theoretically justified and is thoroughly evaluated against the current SVC test model JSVM and DWTSB through extensive coding experiments for scalable coding of intra frame. The simulation results show the proposed scanning algorithm consistently outperforms JSVM and DWTSB in PSNR performance. This results in extra compression for intra frames, along with spatial scalability. Thus Image and video coding applications, traditionally serviced by separate coders, can be efficiently provided by an integrated coding system.
Read moreComputationally efficient and adaptive scalable video coding
Current digital video applications require delivery of digital video content to different clients over heterogeneous networks. These clients may have different system resources and bandwidth capabilities. Thus, delivering of digital video data compressed at specific spatial resolution, quality levels, frame rates and encoding those video data at specific bit rates according to the client resources and the available network is essential. The characteristics of current video transmission system pose many problems for video delivery and scalable video coding is a highly attractive solution. This thesis investigates the existing video coding standards, especially the newest H.264/advanced video coding and its scalable extension: scalable video coding. In this thesis, we address two issues related to the state-of-the-art scalable video coding: 1) computational complexity reduction; and 2) video coding efficiency enhancement. In the first part of this thesis, we address the complexity issue of spatial, signal-to-noise ratio (SNR), temporal and combined scalable video coding. Our goal is to find a good tradeoff between computational complexity and video coding efficiency. We first establish the need to reduce computational complexity for current H.264/advanced video coding and scalable video coding framework. We then evaluate the mode distribution correlations between the base layer and its enhancement layers for different video scalabilities. After the exhaustive search over all possible block partitions is performed at the base layer, the number of candidate modes for luma and chroma blocks in a macroblock that take part in rate distortion optimization calculation at enhancement layers could be greatly reduced based on the correlations. Finally, adaptive fast mode decision schemes for spatial, SNR and temporal scalable video coding are presented. Our schemes could achieve consistent and significant computational complexity reduction with negligible loss in objective and subjective video quality and insignificant increments in bit rate consumption. The second part of this thesis is to provide a solution for the problem of high computational complexity of scalable video decoders. An adaptive encoding algorithm is proposed to reduce decoder complexity in coarse grain SNR scalable video coding.
Read moreOptimised Compression Strategy in Wavelet-Based Video Coding using Improved Context Models
Accurate probability estimation is a key to efficient compression in entropy coding phase of state-of-the-art video coding systems. Probability estimation can be enhanced if contexts in which symbols occur are used during the probability estimation phase. However, these contexts have to be carefully designed in order to avoid negative effects. Methods that use tree structures to model contexts of various syntax elements have been proven efficient in image and video coding. In this paper we use such structure to build optimised contexts for application in scalable wavelet-based video coding. With the proposed approach context are designed separately for intra-coded frames and motion-compensated frames considering varying statistics across different spatio-temporal subbands. Moreover, contexts are separately designed for different bit-planes. Comparison with compression using fixed contexts from embedded ZeroBlock coding (EZBC) has been performed showing improvements when context modelling on tree structures is applied.
Read moreA New Wireless Generation Technology for Video Streaming
With the exponential rise in the volumes of video traffic in cellular networks, there is an urgent need for improving the quality of video delivery. This research proposes a mobile generation model based on the updated technologies of the fourth- and fifth-generation mobile systems, which is called Proposed Generation (Pro-G). This model uses wider bandwidth and advanced adaptive modulation and coding. It also incorporates the method of the adaptive video streaming of multiple video data rates by using the transcoding technique, which is called H.265 proposed (H.265 pro). Thus, both methods are tested to provide a large number of users of video/data application with more speed and best quality. A comparison with 4G technology is done to assign the development regarding number of users with data rate. The suggested video coding shows how much the overall system is more reliable over the congested channel than conventional video coding technologies such as high-efficiency video coding (HEVC/H.265) and advanced video coding (AVC/H.264). The results showed that the proposed method of transmitting wireless data is better than the LTE-ADV method. In this method, the rate of data transfer increases by 29% compared with LTE-ADV, while the bit rate saving was increased to 13% in the proposed video coding compared with that in the H.265.
Read moreMPEG-5 Part 1: Essential Video Coding
<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">The Motion Picture Experts Group (MPEG) standardization group has produced a large number of standards for video compression over the last three decades. Traditionally, the MPEG standards have either focused on highest available compression efficiency [e.g., MPEG-2, advanced video coding (AVC), and high-efficiency video coding (HEVC)] or a desire to produce a royalty-free standard [e.g., Internet video coding (IVC) and web video coding (WebVC)]. In January 2019, MPEG embarked on a new standardization project that can be seen as a hybrid of the two: MPEG-5 part 1 and essential video coding (EVC) [International Organization for Standardization/International Electrotechnical Commission (ISO/IEC) 23094-1]. The EVC standard was developed with a royalty-free baseline profile at its base and a royalty bearing main profile that provides excellent compression performance. The main profile adds, on top of the baseline profile, 21 different coding tools that each can be individually turned off and, when necessary, replaced by a corresponding baseline profile tool. This structure makes it easy to fall back to a smaller set of tools in the future, if, for example, licensing complications occur around a specific tool, without breaking compatibility with already deployed decoders</i> .
Read moreA novel VLSI concurrent dual multiplier-dual adder architecture for image and video coding applications
In this paper, an efficient algorithm for concurrent computation of two real multiplications and/or two real additions usually required for high-throughput image and video coding applications is described. The proposed algorithm is mapped onto a novel concurrent dual multiplier-dual adder cell based on carry-save 4:2 compressors. A detailed performance analysis of the the proposed cell shows reductions ranging from 15% to 60% in the computation time and area when compared with the conventional processing elements making it highly attractive for VLSI implementation.
Read moreAn efficient inter prediction mode selection scheme for advanced video coding based on motion homogeneity and residual complexity
Modern video coding standards such as High Efficiency Video Coding (HEVC) and H.264/MPEG‐4 advanced video coding (AVC) supersede the previous coding standards because of their improved coding efficiency. These standards adopt variable block sizes in frame coding ranging from 4 × 4 to 16 × 16 and 4 × 4 to 64 × 64 for H.264/AVC and HEVC, respectively. The use of variable block sizes for inter prediction provides a significant coding gain compared to coding a macroblock (MB) using regular block size. However, this new feature greatly increases the computational complexity of the encoder when brute‐force rate distortion optimization (RDO) algorithm is used for coding parameter selection. This paper proposes an efficient inter prediction mode selection scheme based on motion homogeneity and residual complexity measures of an MB to speed up the encoding process. The motion homogeneity is assessed through the normalized motion vector (MV) field, and residual complexity is evaluated by the sum of absolute difference (SAD). To acquire the MVs and SADs, motion estimation at 8 × 8 block size is performed using a lightweight recursive motion estimator in which the vector field tends toward true object motion. Based on motion homogeneity and residual complexity of an MB, only a small number of inter prediction modes are selected for the RDO process. The experimental results for H.264/AVC show that the proposed scheme reduces the encoding time by 64% on average without any significant degradation of coding efficiency. © 2016 Institute of Electrical Engineers of Japan. Published by John Wiley & Sons, Inc.
Read moreFlexible Complexity Control Solution for Transform Domain Wyner-Ziv Video Coding
Most Wyner-Ziv (WZ) video coding solutions in the literature focus on improving the coding efficiency. Recently, a few papers have addressed problems related to the complexity distribution of WZ video coding by sharing the motion estimation process between the encoder and decoder. However, these methods turn out to significantly increase the computational complexity of the encoding process, due to the presence of motion estimation at the encoder, which is less appealing in WZ coding applications, since their primary requirement is a low-encoding complexity. To address this problem, we propose a different approach to complexity control based on the adaptive selection between WZ and intra coding in the WZ frames. This solution is more flexible than the previous ones in that it can be implemented either at the encoder or decoder. The proposed intra mode selection algorithm exploits the spatial and temporal coherency in the transform domain Wyner-Ziv video coding (TDWZ), in order to achieve complexity control between the WZ encoder and decoder. The experimental results illustrate that the proposed algorithm not only effectively distributes the computational complexity over the encoder and decoder, but also retains the low-complexity feature at the encoder. This should make it attractive for a large number of real video applications of the WZ video coding paradigm. Moreover, the coding efficiency of the conventional TDWZ codec without intra mode decision is improved by up to 2 dB by the proposed intra mode selection algorithm.
Read moreStudy and investigation of video steganography over uncompressed and compressed domain: a comprehensive review
In the technological era, the primary source of information is in the form of digital data, which has to be secured while storing or transmitting during communication over an unsecured network. Different approaches are used to provide security to digital data, viz. text, audio, image, and video. This paper initially explains the security system such as cryptography, watermarking, and steganography and their comparative analysis based on different characteristics, viz. satisfaction level of objective, type of carrier object and secret information to be used, dependency of security level, and quality assessment parameters. This review article focuses more on steganography methods applied over video. The various methods implemented for video steganography in compressed domain, viz. inter-frame and intra-frame prediction, motion vector estimation, entropy coding (CAVLC and CABAC), and transformed and quantized coefficients of DCT, DST, and DWT, etc. and the methods based on spatial and transform domain for uncompressed video are briefly described. It is followed by the detailed analysis of related work done by various researchers in video steganography and the obtained experimental results. Furthermore, the confidential data hiding in compressed videos are explained using Moving Picture Expert Group (MPEG—1, MPEG—2, MPEG—4), Advanced Video Coding (AVC)/H.264, and High-Efficiency Video Coding (HEVC)/H.265 that includes both spatial and transform domain. This paper summarizes and explains the detailed investigations of numerous techniques of video steganography based on the comprehensive literature survey. The methods used to assess the performance of video steganography are analyzed based on the quality assessment parameters such as imperceptibility; measured by peak signal to noise ratio (PSNR), mean square error (MSE), and structural similarity (SSIM), robustness; measured by bit error rate (BER) and similarity (Sim), and embedding capacity; measured by hiding ratio. The overall review of past literature facilitates to have in-depth knowledge for upgrading the video steganography.
Read more