Speaking Browser: On Patterns in Web Browsing History and the Efficacy of Seq2Seq Models for Access Prediction

Loading...
Thumbnail Image

Files

SHERMAN-JUSTIN-THESIS.pdf (3.12 MB)

Date

2024-07-13

Journal Title

Journal ISSN

Volume Title

Publisher

Access Restrictions

Walk-in Access. This thesis can only be viewed on computer terminals at the <a href=http://mudd.princeton.edu>Mudd Manuscript Library</a>.

Abstract

This thesis investigates a dataset of 11 months of my personal browsing history. It tests the viability of LSTM and transformer-based models for predicting the websites a user will access next given their browsing history. The approach divides browser history entries into browsing session URL groups that are further divided into source and target URL sequences. Embeddings are learned for unique page URLs and, in one trial, month-time slots. Exploration of the dataset suggests some predictability, but while models show promise in learning training data, they fail to predict future accesses for most held-out testing data. Successful predictions generated from test data often include the most frequently and consistently accessed pages, like Gmail.

Description

item.page.type

Princeton University Senior Theses

Keywords

item.page.location

Citation