Using a Power Law Distribution to describe Big Data
arXiv:1509.00504 · doi:10.1109/HPEC.2015.7322459
Abstract
The gap between data production and user ability to access, compute and produce meaningful results calls for tools that address the challenges associated with big data volume, velocity and variety. One of the key hurdles is the inability to methodically remove expected or uninteresting elements from large data sets. This difficulty often wastes valuable researcher and computational time by expending resources on uninteresting parts of data. Social sensors, or sensors which produce data based on human activity, such as Wikipedia, Twitter, and Facebook have an underlying structure which can be thought of as having a Power Law distribution. Such a distribution implies that few nodes generate large amounts of data. In this article, we propose a technique to take an arbitrary dataset and compute a power law distributed background model that bases its parameters on observed statistics. This model can be used to determine the suitability of using a power law or automatically identify high degree nodes for filtering and can be scaled to work with big data.
5 pages
References in corpus (4)
Cited by in corpus (15)
- Static Graph Challenge: Subgraph Isomorphism
- LaraDB: A Minimalist Kernel for Linear and Relational Algebra Computation
- GraphChallenge.org: Raising the Bar on Graph Analytic Performance
- PageRank Pipeline Benchmark: Proposal for a Holistic System Benchmark for Big-Data Platforms
- From NoSQL Accumulo to NewSQL Graphulo: Design and Utility of Graph Algorithms inside a BigTable Database
- Design, Generation, and Validation of Extreme Scale Power-Law Graphs
- GraphChallenge.org Sparse Deep Neural Network Performance
- GraphChallenge.org Triangle Counting Performance
- Distributed Triangle Counting in the Graphulo Matrix Math Library
- AI Enabling Technologies: A Survey
- Benchmarking the Graphulo Processing Framework
- Technical Report on Data Integration and Preparation
- Database Operations in D4M.jl
- D4M 3.0: Extended Database and Language Capabilities
- Technical Report: Developing a Working Data Hub