Publication:

From Flattened Directories to Research Data An LLM-Assisted Pipeline for Structuring Seventh-day Adventist Yearbooks, 1883-1920

Loading...
Thumbnail Image

Files

written_final_report.pdf (1.21 MB)

Date

2026-04-13

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

This thesis presents a workflow, built around a large language model, for transforming scanned Seventh-day Adventist Yearbooks into structured historical data. The Yearbooks are annual administrative directories that record officers, ministers, licentiates, missionaries, institutions, conferences, and organizational units. They are valuable because they preserve the denomination’s own administrative categories during a period when the church was centralizing rapidly, expanding geographically, and layering new bureaucratic structures on top of old ones. They are difficult to interpret because they survive primarily as scanned PDFs whose OCR output flattens visual hierarchy. The problem is both name recognition, creating and cleaning a dataset and figuring out which conference, mission, institution, or department each person should be assigned to once the typography that carried that meaning on the printed page has been lost.

Early experiments in rule-based parsing alone weren't able to produce meaningful results. These deterministic parsing experiments helped diagnose the source’s structure, but they broke down on the core requirement of preserving hierarchy. The project described here rests fully on the LLM parsing implementation of this project. A language model performs the primary record extraction followed by deterministic routines that handle normalization, cleaning, aggregation, and interface logic afterwards. The thesis documents this workflow end to end: corpus preparation, prefiltering, prompt design, JSON extraction, rate limiting, caching, normalization, and browser-based exploratory analysis. Using the currently exported corpus of twenty-eight CSV files covering 1883-1894 and 1904-1920, the project assembles 46,109 row-level records and 9,750 distinct names. The extracted dataset records 227 conference labels, 805 organization labels, and 206 institution labels. It reveals strong evidence of institutional growth and title differentiation. Extracted rows rise from 858 in 1883 to 3,186 in 1894, while conference-label diversity also grows sharply across the same period. The resumption of the Yearbook in 1904 occurs in a reorganized denominational environment and registers a church whose reporting habits are more organized as a whole. Women’s participation is also visible in the corpus. Among uniquely named individuals with inferable gender, women rise from 13.3% in 1883 to 31.6% in 1920, with especially strong representation in missionary licentiate, Bible-work, secretarial, educational, and secretary-treasurer roles. The thesis treats the Yearbook as a longitudinal source for studying institutional growth, administrative geography, and the gendered distribution of titled work. In contrast to how the Yearbook might have been used in the denomination, this thesis attempts to aggregate information about Seventh Day Adventist leaders across the years as opposed to merely within one year. This thesis argues that semi-structured historical directories demand a layered computational strategy, extraction using LLMs for the complex structure of the historical document, deterministic normalization for cross-document consistency, and finally visualization to illuminate and quantify historical claims about the development of the Church.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation