Detecting Memory and Structure in Human Navigation Patterns Using Markov Chain Models of Varying Order
arXiv:1402.0790 · doi:10.1371/journal.pone.0102070
Abstract
One of the most frequently used models for understanding human navigation on the Web is the Markov chain model, where Web pages are represented as states and hyperlinks as probabilities of navigating from one page to another. Predominantly, human navigation on the Web has been thought to satisfy the memoryless Markov property stating that the next page a user visits only depends on her current page and not on previously visited ones. This idea has found its way in numerous applications such as Google's PageRank algorithm and others. Recently, new studies suggested that human navigation may better be modeled using higher order Markov chain models, i.e., the next page depends on a longer history of past clicks. Yet, this finding is preliminary and does not account for the higher complexity of higher order Markov chain models which is why the memoryless model is still widely used. In this work we thoroughly present a diverse array of advanced inference methods for determining the appropriate Markov chain order. We highlight strengths and weaknesses of each method and apply them for investigating memory and structure of human navigation on the Web. Our experiments reveal that the complexity of higher order models grows faster than their utility, and thus we confirm that the memoryless model represents a quite practical model for human navigation on a page level. However, when we expand our analysis to a topical level, where we abstract away from specific page transitions to transitions between topics, we find that the memoryless assumption is violated and specific regularities can be observed. We report results from experiments with two types of navigational datasets (goal-oriented vs. free form) and observe interesting structural differences that make a strong argument for more contextual studies of human navigation in future work.
References in corpus (2)
Cited by in corpus (27)
- Representing higher-order dependencies in networks
- Why We Read Wikipedia
- HypTrails: A Bayesian Approach for Comparing Hypotheses About Human Trails on the Web
- Effects of memory on spreading processes in non-Markovian temporal networks
- What Makes a Link Successful on Wikipedia?
- Discovering Beaten Paths in Collaborative Ontology-Engineering Projects using Markov Chains
- Predictive Bayesian selection of multistep Markov chains, applied to the detection of the hot hand and other statistical dependencies in free throws
- Dynamics and Biases of Online Attention: The Case of Aircraft Crashes
- Maps of sparse Markov chains efficiently reveal community structure in network flows with memory
- Entropy estimators for Markovian sequences: A comparative analysis
- How to Apply Markov Chains for Modeling Sequential Edit Patterns in Collaborative Ontology-Engineering Projects
- Wikipedia Reader Navigation: When Synthetic Data Is Enough
- Statistical analysis of the first passage path ensemble of jump processes
- MixedTrails: Bayesian hypothesis comparison on heterogeneous sequential data
- Query for Architecture, Click through Military: Comparing the Roles of Search and Navigation on Wikipedia
- Predicting Sequences of Traversed Nodes in Graphs using Network Models with Multiple Higher Orders
- Why the World Reads Wikipedia: Beyond English Speakers
- Random Surfers on a Web Encyclopedia
- How PHP Releases Are Adopted in the Wild?
- Control Flow Graph Modifications for Improved RF-Based Processor Tracking Performance
- HopRank: How Semantic Structure Influences Teleportation in PageRank (A Case Study on BioPortal)
- Individual differences in knowledge network navigation
- Learning the Markov order of paths in a network
- Information-theoretic analysis of temporal dependence in discrete stochastic processes: Application to precipitation predictability
- An improved estimator of Shannon entropy with applications to systems with memory
- Broccoli: Sprinkling Lightweight Vocabulary Learning into Everyday Information Diets
- Consensus dynamics on temporal hypergraphs