• Home
  • Search
  • Exploiting Data-pattern-aware Vertical Partitioning to Achieve Fast and Low-cost Cloud Log Storage
  • Cite Icon1
  • https://doi.org/10.1145/3643641Copy DOI Icon

Exploiting Data-pattern-aware Vertical Partitioning to Achieve Fast and Low-cost Cloud Log Storage

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Cloud logs can be categorized into on-line, off-line, and near-line logs based on the access frequency. Among them, near-line logs are mainly used for debugging, which means they prefer a low query latency for better user experience. Besides, the storage system for near-line logs prefers a low overall cost including the storage cost to store compressed logs, and the computation cost to compress logs and execute queries. These requirements pose challenges to achieving fast and cheap cloud log storage. This article proposes LogGrep, the first log compression and query tool that exploits both static and runtime patterns to properly structurize and organize log data in fine-grained units. The key idea of LogGrep is “vertical partitioning”: it stores each log entry into multiple partitions by first parsing logs into variable vectors according to static patterns and then extracting runtime pattern(s) automatically within each variable vector. Based on such runtime patterns, LogGrep further decomposes the variable vectors into fine-grained units called “Capsules” and stamps each Capsule with a summary of its values. During the query process, LogGrep can avoid decompressing and scanning Capsules that cannot match the keywords, with the help of the extracted runtime patterns and the Capsule stamps. We further show that the interactive debugging can well utilize the advantages of the vertical-partitioning-based method and mitigate its weaknesses as well. To this end, LogGrep integrates incremental locating and partial reconstruction to mitigate the read amplification incurred by vertical-partitioning-based method. We evaluate LogGrep on 37 cloud logs from the production environment of Alibaba Cloud and the public datasets. The results show that LogGrep can reduce the query latency and the overall cost by an order of magnitude compared with state-of-the-art works. Such results have confirmed that it is worthwhile applying a more sophisticated vertical-partitioning-based method to accelerate queries on compressed cloud logs.

Similar Papers
  • Research Article

A clique-based scheduling in real-time query processing optimisation for cloud-based wireless body area networks

  • Jan 01, 2019
  • International Journal of Biomedical Engineering and Technology
  • Dinesh Kumar A +1
  • Conference Article
  • Citations26

Declarative Network Monitoring with an Underprovisioned Query Processor

  • Jan 01, 2006
  • F Reiss +1
  • Research Article
  • Citations2

Spoken conversational search: Evaluating the effect of system clarifications on user experience through Wizard‐of‐Oz study

  • Dec 29, 2024
  • Journal of the Association for Information Science and Technology
  • Souvick Ghosh +1
  • Book Chapter
  • Citations8

MOSS-DB: A Hardware-Aware OLAP Database

  • Jan 01, 2010
  • Yansong Zhang +2
  • PDF
  • Research Article
  • Citations63

Convolutional Neural Network Based Classification of App Reviews

  • Jan 01, 2020
  • IEEE Access
  • Naila Aslam +3
  • Research Article

Robust 2.5D Feature Matching in Light Fields via a Learnable Parameterized Depth-Degraded Projection.

  • Jan 01, 2026
  • IEEE transactions on image processing : a publication of the IEEE Signal Processing Society
  • Meng Zhang +4
  • Research Article
  • Citations3

Energy-Efficient Resource Allocation for HARQ With Statistical CSI

  • Dec 01, 2018
  • IEEE Transactions on Vehicular Technology
  • Xavier Leturc +2
  • Research Article
  • Citations1

Enhancing GP consultation skills training: educational evaluation of a conversational AI innovation for simulated consultation assessment preparation

  • Sep 14, 2025
  • Education for Primary Care
  • Chris Jacobs +2
  • Book Chapter

Bitmap Join Indexes vs. Data Partitioning

  • Jan 01, 2009
  • Ladjel Bellatreche
  • Research Article
  • Citations19

Venus: Scalable Real-Time Spatial Queries on Microblogs with Adaptive Load Shedding

  • Feb 01, 2016
  • IEEE Transactions on Knowledge and Data Engineering
  • Amr Magdy +4
  • Book Chapter
  • Citations1

Learning-Based Optimization for Online Approximate Query Processing

  • Jan 01, 2022
  • Wenyuan Bi +5
  • Research Article

Semantic Data Lakes: Integrating Big Data and Knowledge Graphs for Enterprise Decision Support

  • Aug 23, 2025
  • International Journal on Advanced Electrical and Computer Engineering
  • Sathish Kaniganahalli Ramareddy
  • Book Chapter
  • Citations1

Use of Self-Service Query Tools Varies by Experience and Research Knowledge

  • Jan 01, 2015
  • Hruby Gregory W +2
  • Research Article
  • Citations3

A Range Query Method for Data Access Pattern Protection Based on Uniform Access Frequency Distribution

  • Jan 01, 2023
  • Journal of Networking and Network Applications
  • Jing Yan +3
  • Book Chapter

Faster MaxScore Document Retrieval with Aggressive Processing

  • Jan 01, 2014
  • Kun Jiang +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.