• Cite Icon9
  • https://doi.org/10.14778/3007263.3007266Copy DOI Icon

Kodiak

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Turn's online advertising campaigns produce petabytes of data. This data is composed of trillions of events, e.g. impressions, clicks, etc., spanning multiple years. In addition to a timestamp, each event includes hundreds of fields describing the user's attributes, campaign's attributes, attributes of where the ad was served, etc. Advertisers need advanced analytics to monitor their running campaigns' performance, as well as to optimize future campaigns. This involves slicing and dicing the data over tens of dimensions over arbitrary time ranges. Many of these queries need to power the web portal to provide reports and dashboards. For an interactive response time, they have to have tens of milliseconds latency. At Turn's scale of operations, no existing system was able to deliver this performance in a cost effective manner. Kodiak, a distributed analytical data platform for web-scale high-dimensional data, was built to serve this need. It relies on pre-computations to materialize thousands of views to serve these advanced queries. These views are partitioned and replicated across Kodiak's storage nodes for scalability and reliability. They are system maintained as new events arrive. At query time, the system auto-selects the most suitable view to serve each query. Kodiak has been used in production for over a year. It hosts 2490 views for over three petabytes of raw data serving over 200K queries daily. It has median and 99% query latencies of 8 ms and 252 ms respectively. Our experiments show that its query latency is 3 orders of magnitude faster than leading big data platforms on head-to-head comparisons using Turn's query workload. Moreover, Kodiak uses 4 orders of magnitude less resources to run the same workload.

Similar Papers
  • Conference Article
  • Citations8

RFEED: A Mixed Workload Scheduler for Enterprise Data Warehouses

  • Mar 01, 2009
  • Proceedings - International Conference on Data Engineering
  • Abhay Mehta +3
  • Research Article
  • Citations1

SiriusBI: A Comprehensive LLM-Powered Solution for Data Analytics in Business Intelligence

  • Aug 01, 2025
  • Proceedings of the VLDB Endowment
  • Jie Jiang +13
  • Research Article

Abstract 469: A multi-scale analysis and visualization platform for cancer data - deriving tumor microenvironment behavior from pathology and transcriptomics

  • Jun 15, 2022
  • Cancer Research
  • Joseph R Peterson +5
  • Conference Article
  • Citations120

Focus: querying large video datasets with low latency and low cost

  • Jan 12, 2018
  • Kevin Hsieh +7
  • Conference Article
  • Citations84

TILDE: Term Independent Likelihood moDEl for Passage Re-ranking

  • Jul 11, 2021
  • Shengyao Zhuang +1
  • Conference Article
  • Citations8

Cloud Native Data Platform for Network Telemetry and Analytics

  • Oct 25, 2021
  • Daniel Tovarnak +2
  • Research Article
  • Citations19

Establishing the Advanced Disaster Reduction Management System by Fusion of Real-Time Disaster Simulation and Big Data Assimilation

  • Mar 01, 2016
  • Journal of Disaster Research
  • Shunichi Koshimura
  • Research Article
  • Citations16

Modelling and analysis of big data platform group adoption behaviour based on social network analysis

  • Mar 29, 2021
  • Technology in Society
  • Zhimei Lei +2
  • Research Article
  • Citations6

Efficient evaluation of generalized tree-pattern queries on XML streams

  • Apr 28, 2010
  • The VLDB Journal
  • Xiaoying Wu +2
  • Conference Article
  • Citations15

SPDO: High-throughput road distance computations on Spark using Distance Oracles

  • May 01, 2016
  • Shangfu Peng +2
  • Research Article
  • Citations4

EAR-Oracle: On Efficient Indexing for Distance Queries between Arbitrary Points on Terrain Surface

  • May 26, 2023
  • Proceedings of the ACM on Management of Data
  • Bo Huang +3
  • PDF
  • Book Chapter
  • Citations12

Findere: Fast and Precise Approximate Membership Query

  • Jan 01, 2021
  • Lucas Robidou +1
  • PDF
  • Conference Article
  • Citations11

A Scalable Index for Top-k Subtree Similarity Queries

  • Jun 25, 2019
  • Daniel Kocher +1
  • Research Article
  • Citations29

On temporal-constrained sub-trajectory cluster analysis

  • Apr 04, 2017
  • Data Mining and Knowledge Discovery
  • Nikos Pelekis +4
  • Research Article
  • Citations64

Occurrence and Spatiotemporal Dynamics of Pharmaceuticals in a Temperate-Region Wastewater Effluent-Dominated Stream: Variable Inputs and Differential Attenuation Yield Evolving Complex Exposure Mixtures.

  • Sep 22, 2020
  • Environmental Science & Technology
  • Hui Zhi +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.