Operations Research and Financial Engineering, 2000-2026

Permanent URI for this collectionhttps://theses-dissertations.princeton.edu/handle/88435/dsp011r66j119j

Browse

Recent Submissions

Now showing 1 - 20 of 129
  • MARKET INFORMATION FROM PREDICTION MARKET AND ITS POTENTIAL AS A FINANCIAL INSTRUMENT

    (2026-04-09) Zhao, Grace; Cattaneo, Matias Damian

    Prediction markets have grown rapidly as platforms for aggregating beliefs about future events, yet their role as financial instruments remains not well understood, particularly relative to traditional markets. This is especially relevant for macroeconomic outcomes such as Federal Reserve interest rate decisions and gold prices, where parallel expectation measures already exist in established financial instruments.

    This paper studies prediction markets as state-contingent financial assets and evaluates their behavior alongside traditional benchmarks. Using data from Polymarket, CME FedWatch, U.S. Treasury bills, rate-sensitive indices, and gold futures, the analysis compares belief dynamics through time-series methods and examines their economic relevance through portfolio construction. Prediction market positions are incorporated as overlays within traditional portfolios to assess their impact on risk and return characteristics, with additional analysis of how contract structure, such as discretized outcome bins, influences pricing and information representation.

    The results show that prediction markets are becoming increasingly aligned with market-implied probabilities, particularly for central outcomes, while exhibiting meaningful variation across event windows that reflects evolving macroeconomic conditions. They exhibit varying volatility, VaR, ES, and other financial characteristics, making them a plausible candidate for hedging. Portfolio experiments indicate that overlays concentrated in high-probability outcomes can improve risk-adjusted performance, whereas tail-outcome positions tend to underperform due to their binary payoff structure. Despite their simplicity, these portfolio constructions demonstrate that prediction market signals can contribute to diversification. Overall, prediction markets aggregate economically relevant information and offer a distinct signal complementary to traditional instruments, though their effectiveness remains dynamic and dependent on market structure and external conditions.

  • Identifying Relative-Value Opportunities Across Investment Grade Corporate Bond Issuers

    (2026-04-14) Liu, Daniel; Scheinerman, Daniel

    This thesis investigates whether investment grade corporate bonds issued by companies in different industries exhibit stable pricing relationships that can create relative-value opportunities. The paper uses TRACE bond trade data and Treasury benchmark yields from 2024 to 2025 to construct a dataset of daily corporate bond spreads. This dataset is then used in the multi-stage empirical model. The first stage is the Johansen cointegration and vector error-correction methods to identify long-run equilibrium relationships. A mean-reversion and half life behavior is then used. The final steps are a two-state regime-switching model and a network mapping model. The results show that there are cross-sector pricing relationships, but these relationships are not universal. Only a meaningful subset of bond pairs that is both economically relevant and statistically compelling display these pricing relationships. The strongest bond pairs form a broader issuer connected network. These findings suggest that cross-sector bond pairs with a stable long-run pricing relationship exist in the investment grade corporate bond market and that these cross-sector relationships are consistent with broader market forces, such as funding conditions and limits to arbitrage.

  • The Weight of Gold: Modeling Cultural Preferences in Optimal Portfolio Allocation

    (2026-04-13) Tanti, Prapti R.; Holen, Margaret

    This thesis studies how a culturally meaningful asset such as gold can enter an optimal portfolio, and which type of portfolio model gives the most convincing explanation for that behavior. The project is motivated by a gap between real household behavior and standard financial models. In many immigrant and culturally rooted households, gold is not viewed only as a speculative asset. It can also function as precautionary savings, family wealth, social status, and a store of value tied to trust, tradition, and financial experience. Standard portfolio theory often leaves little room for that kind of demand. To study this question, we begin with two financial benchmarks, a mean–variance model and a goals-based model, and then extend the goals-based framework so that cultural preference for gold can enter in three different ways. The results show that gold does not enter for the same reason in every model. In the mean–variance benchmark, gold can appear because its lower volatility and diversification value improve the portfolio trade-off. In the goals-based benchmark, gold enters only in narrower cases where it helps raise the probability of reaching a target. Among the cultural extensions, a simple linear preference term becomes too unstable, while threshold and penalty formulations generate more realistic moderate gold demand. The penalty specification performs best overall because it allows a bounded cultural role for gold while preserving a meaningful financial trade-off. A descriptive empirical bridge using the SCF and SIPP datasets supports a moderate and proxy-dependent cultural-demand story rather than either zero shift or extreme gold preference.

  • Coupling LLMs with KGs: Benchmarking LLM Models and Structured Reasoning Frameworks with a Classic Botanical Field Guide

    (2026-04-13) Matters, John; Holen, Margaret

    Large language models remain susceptible to hallucination and stochastic outputs when reasoning over structured domains. This thesis investigates whether coupling vision-language models with a typed knowledge graph can produce verifiable botanical identification, and whether ensemble disagreement across model runs can serve as a scalable signal for detecting prediction uncertainty without human annotation. The experimental platform is Newcomb’s Wildflower Guide, a hierarchical botan- ical key, whose feature-value pairs are encoded as a knowledge graph and paired with a corpus of expression-unlabeled field photographs from iNaturalist. A two-stage mapping pipeline, first an existence check followed by a blind multiple-choice clas- sification, is applied across six vision-language models. A human-curated reference set establishes perceptual alignment between models and the Newcomb vocabulary before evaluation on unlabeled images. The central finding is that ensemble disagreement is a reliable, annotation-free sig- nal for image-level uncertainty. Critically, same-model repetition for the most stable model produces sparser but more precise disagreement than architecturally diverse ensembles: when a stable model disagrees with itself on a specific observation, the disagreement is more concentrated on true errors, whereas cross-model disagreement conflates image-level ambiguity with model architectural bias. A further finding is that unanimous disagreement between models and the Newcomb key can surface candidate annotation errors at scale without any additional human effort. Together, these results establish a practical path for extending the pipeline to large corpora of expression-unlabeled images and for generalizing the methodology to domains with similarly structured knowledge.

  • Detecting, Visualizing, and Trading on Dynamic Thematic Structure in Equity Markets

    (2026-04-12) Henriques, Samuel H.; Almgren, Robert

    We develop a systematic framework for detecting and visualizing dynamic equity themes, time-varying clusters of stocks experiencing excess co-movement beyond standard risk factors, and study their implications for options markets and trading. Using a fixed universe of liquid US stocks, daily returns are first residualized with respect to the Fama-French five factors and momentum to produce return series that isolate dependence beyond broad market exposures, additionally residualized with respect to sector returns for a robustness check. Rolling correlation networks built from these residuals are then analyzed for thematic structure as well as emergence, shocks, and dissipation, using HDBSCAN for clustering and UMAP with Procrustes alignment to produce coherent visualizations of evolving geometry. Resulting clusters are linked into multi-day lifecycle objects, then we conduct episode-level analyses and state-based modeling and trading tests using options-derived variables. This framework defines themes from residualized dependence rather than from narratives or static industry classifications and treats options markets as a separate layer for mechanism analysis and trading applications rather than as a detection input. Empirically, we find that detected themes exhibit repeating options signatures over their lifecycles, and we identify a theme-conditioned options trading strategy that generates positive payoff, suggesting a potential application of the framework in systematic trading and risk management.

  • The “Moneyball” Myth: Evaluating Public Pre-Draft Data Against Expert Scouting in the NFL

    (2026-04-09) Brandl, Ryan C.; Rigobon, Daniel

    The National Football League (NFL) Draft is a high-stakes selection process where teams commit multi-million dollar investments to unproven talent under high uncertainty of future production. “Moneyball” methodology transformed baseball and sports analytics as a whole, offering the potential to capitalize on inefficient traditional approaches to player valuation. This thesis tests whether a similar approach can improve NFL Draft decisions. To do this, a comprehensive feature set with in-depth college statistics, combine results, and additional physical and school-context measurements is designed. Linear and nonlinear models are trained to predict Approximate Value, a position-independent output metric of on-field contribution, across various time horizons. The predictive power of this set is evaluated against ESPN scouting variables, which are aggregated expert grades and rankings of incoming prospects. Predicted target values are utilized to test player selection in historical drafts and reconstruct traditional pick-value curves to better estimate drafting capital. Results showed that without supplemental scouting information, public data alone could not predict performance on any time horizon. Furthermore, a simple calibration of a single ESPN scouting variable outperformed the full feature set machine learning pipeline across all horizons. Reranking historical drafts through model predictions achieved modest results. This implies that scout-level knowledge, although subjective, already absorbs the signal in publicly available data.

  • Latency Arbitrage Across Centralized and Decentralized Exchanges: Measuring Gap Survival on L1 vs L2

    (2025-10-10) Pandit, Arnav; Hanin, Boris

    Does reducing the block confirmation delay to 48 times faster, down from 12 seconds on Ethereum mainnet (L1) to 0.25 seconds on Arbitrum One (L2), create a significant efficiency gain for DEXs from a latency arbitrage perspective? To answer this question, four weeks of minute-by-minute data between February 19 and March 17, 2026, are used to estimate the CEX-DEX price spread for ETH and WBTC pools on Uniswap v3 against Binance US, KuCoin, Coinbase, and Bitstamp. The τ ∗ criterion establishes the profitability threshold separating exploitable price gaps from false positives, while the methodology covers a total of five analytical steps: gap distributions, Kaplan-Meier survival analysis, and 1,000 replications of a Monte Carlo backtest, and four robustness checks. L2 reduces the mean ETH CEX-DEX spread by 70% (from 69.17 to 21.01 basis points) and the upper bound of the 90% confidence interval of realized P&L by 54% for ETH and 78–82% for WBTC, suggesting that L2 is an execution efficiency upgrade rather than price discovery technology. For WBTC, where the 0.30% pool fee imposes a lower bound of τ ∗ = 95.70 bps, 98.5% of gap-minutes on L2 will be unprofitable to exploit even at a trade size of $10,000. The improvement in variance increases under volatility stress and grows from 51% in normal conditions to 76% during high-volatility regimes. The key take-away from this research can be summarized into the following hierarchy of frictions: CEX fee structure > chain architecture > pool fee > block time > gas cost. For DEX developers, this finding means that switching a low-fee pool to L2 decreases the LPs’ expected loss-to-rebalancing cost and increases execution efficiency for arbitrageurs. For high-fee pools, switching to L2 improves execution quality but does not improve LPs’ welfare.

  • Growing the Game, Shrinking the Player: The Structural Failure of Professional Squash, 2018–2025

    (2026-04-09) Khalil, Hassan; Rigobon, Daniel

    Professional squash has experienced substantial prize money growth between 2018 and 2025, yet the economic conditions facing the majority of PSA World Tour players have stagnated or worsened. Squash’s inclusion in the 2028 Los Angeles Olympics has been celebrated as validation of the sport’s trajectory. The data suggest a different reading: Olympic inclusion is arriving at precisely the moment when the economic infrastructure required to capitalise on it does not exist. This thesis develops a nine-indicator Economic Welfare Score (EWS), comprising inequality, scale, and player welfare dimensions, and applies it to PSA tournament and player earnings data across seven seasons (2018–2025). The analysis establishes that PSA’s composite economic score in 2024–25 remains below its 2018–19 baseline. Benchmarked against the ATP (Association of Tennis Professionals) across all nine indicators, the divergence is stark: ATP’s equivalent score has risen substantially over the same period, driven by institutional capital, broadcast infrastructure, and a restricted player pool.

    Five econometric models confirm the structural mechanisms driving this divergence. Rank is a near-deterministic predictor of whether a player can earn a living wage, independent of season-level conditions. Calendar expansion has mechanically worsened earnings inequality, as tournament growth has outpaced prize pool growth and compressed the median purse. These relationships are stable across all model specifications.

    The thesis concludes that PSA prize pool growth represents structural divergence from ATP, not convergence. The tour has expanded its roster considerably, but the overwhelming majority of additional players earn below any reasonable subsistence threshold. Structural reform, comprising centralised broadcast rights, a restricted player ceiling, and prize redistribution away from lower-tier events, is the prerequisite for economic progress, not further calendar expansion.

  • Leveraging Decision-Focused Learning to Optimize Stroke Patients' CT Scanning Regimens

    (2026-04-09) Yaninek, Zachary J.; Stellato, Bartolomeo

    The United States’ healthcare sector currently faces a variety of challenges, including rising costs, overcrowding, a lack of personalized medical care, and physician burnout. Incorporating optimization, machine learning, and artificial intelligence within healthcare systems (often referred to as healthcare analytics) has the potential to address many of these shortcomings. We apply machine learning and optimization to the process of sequentially planning computed tomography (CT) scans across multiple patients within a neurology intensive care unit. We model this as a rolling horizon value maximization integer program where the value of scanning each patient at each time step within the planning horizon is uncertain and must be predicted from data. We evaluate the model’s performance on the basis of its event alignment (i.e., if it schedules patient scans appropriately). Motivated to predict these scanning values such that they produce optimal scheduling decisions, we construct a neural network regressor we term RewardNet and train it using the decision-focused learning (DFL) SPO+ loss function. We found that the DFL approach consistently outperforms other regressor loss functions and a multi-class classifier baseline in both offline and online simulation-driven out-of-sample performance tests and admits an interpretable feature importance structure. Nevertheless, although the rolling horizon oracle drastically outperforms the historical clinical decision baseline, our DFL-driven model falls short of this metric. This motivates further research to develop a DFL-driven model that can surpass the clinical baseline and eventually serve as a clinical decision support tool.

  • Risk-Adjusted Valuation of Agentic-AI Firms: Hierarchical Cohort Unit Economics, Reliability Learning, Conformal Forecasting, and ES-Based Downside Mapping

    (2026-04-09) Gupta, Hazel; Scheinerman, Daniel

    This thesis examines whether agentic-AI firms can be valued adequately using the roll-up discounted cash flow frameworks commonly applied to SaaS businesses. It argues that such firms present valuation challenges that are not well captured by conventional approaches based on blended churn assumptions, smooth margin expansion, and discount-rate-based risk adjustment alone. Because agentic-AI businesses frequently exhibit usage-linked monetization, reliability-sensitive retention, and asymmetric operating downside, their economics are more naturally understood through a framework that models forecast uncertainty and tail risk explicitly. To study this problem, the thesis develops a risk-adjusted valuation framework comprising four components: cohort-level unit economics, a time-varying reliability signal, conformal predictive intervals, and downside valuation through Expected Shortfall. The framework is implemented in a synthetic environment designed to generate cohort heterogeneity, reliability learning, regression shocks, reliability-linked tail events, and evolving cost structures. Forecasts and valuations are evaluated under a strict time-cut protocol that excludes look-ahead bias. The empirical pipeline is fully reproducible and generates the synthetic panel, roll-up series, forecast intervals, valuation distributions, robustness results, tables, and figures. The empirical findings are mixed but informative. In the current implementation, the benchmark roll-up model attains stronger empirical coverage at the nominal 80% and 90% levels, whereas the reliability-aware framework produces slightly narrower intervals and materially different left-tail valuation statistics. These results indicate that explicit conditioning on reliability materially affects downside valuation even before the full hierarchical model is estimated. At the same time, the remaining valuation error indicates that improved state estimation and model-consistent uncertainty calibration are necessary before the proposed framework can be regarded as superior to a disciplined roll-up benchmark. The principal contribution of the thesis is therefore methodological rather than dispositive. It establishes a testable end-to-end framework for calibrated forecasting and tail-aware valuation of agentic-AI firms, shows that explicit reliability modeling materially changes downside valuation, and identifies the precise components, most notably latent-state estimation and model-consistent uncertainty calibration, that must be strengthened before broader claims of superiority can be sustained.

  • Bidding Equilibria in Electricity Markets: ERCOT, PJM, and the Mean Field Limit

    (2026-05-09) Yeung, Sabrina; Tangpi, Ludovic

    As electricity markets grow larger and more uncertain with increasing renewable penetration, a central question is whether standard equilibrium models of strategic bidding remain valid at scale. This thesis evaluates that question in two steps. First, I apply a static supply-function equilibrium (SFE) framework to two U.S. wholesale electricity markets with distinct designs: ERCOT and PJM. In ERCOT, the model overpredicts strategic withholding but captures the qualitative direction of price movements. In PJM, the model breaks down more fundamentally, producing nearly load-invariant equilibrium prices despite substantial variation in observed outcomes. Second, I examine the computational limitations of finite-player Nash equilibrium. Scaling experiments show that best-response algorithms fail to converge beyond approximately N ≈ 20 firms, making the approach infeasible at realistic market sizes. To address this, I implement a mean field game (MFG) approximation, which replaces the high-dimensional fixed-point problem with a representative agent formulation. Calibrated to PJM data, the mean field equilibrium closely matches the finite-player Nash solution, consistent with theoretical predictions. However, both models fail to reproduce observed price dynamics in similar ways. This suggests that discrepancies between model predictions and real market outcomes arise primarily from omitted operational features, such as transmission constraints, unit commitment, and intertemporal coupling, rather than from the modeling of strategic interaction itself.

  • Learning-Guided Humanitarian Facility Allocation: An Integrated Decision-Support System for Global Emergency Response

    (2026-05-09) Gaherwar, Samiksha; Dytso, Alex

    This thesis presents a three-module decision-support system for the first operational window after a humanitarian emergency alert on the IFRC GO platform, the International Federation of Red Cross and Red Crescent Societies' operational data environment for events, field reports, and facilities. Service needs are estimated as multi-label probabilities from historical field reporting and event features; those estimates then feed two downstream components—a hybrid retriever (sparse lexical scoring plus dense semantic similarity) that queries operational lessons using hazard, geography, and predicted sectors, and a mixed-integer assignment model that maps thresholded demand to geocoded local units under feasibility, capacity, and distance- and border-related penalties. External data include EM-DAT (Emergency Events Database) impact fields and INFORM country-year risk indices to give scale and structural-risk context beyond IFRC's own severity scores.

    On a temporal test split at January 1, 2024, tree-based and linear estimators improve on majority and large language model baselines for service prediction; hybrid retrieval achieves the strongest pilot ranking scores among the configurations compared, with time-decay and cross-encoder ablations documented in the results chapter; and the assignment model remains feasible across six evaluation scenarios while flagging services that lack a capable facility within range. The discussion interprets when heuristic and optimal assignments coincide, how retrieval design constrains language-model re-ranking, and what facility metadata would sharpen optimization in future work.

  • Data Centers and Electricity Prices in PJM: An Analysis of Load Growth and Trading Opportunities

    (2026-04-09) Stone, William B.; Dytso, Alex

    Northern Virginia is home to the world’s largest concentration of data centers, with over 900 facilities and more than 89,000 megawatts of capacity all in a twenty mile radius in Loudoun County. These buildings operate at an almost constant load around the clock, creating persistent and concentrated demand on a system not designed to accommodate it. Despite the huge growth of data centers recently, no prior empirical work has quantified their causal effect on local electricity congestion prices or if it opens trading opportunities. This thesis asks two questions. First, does the concentration of data center load in Northern Virginia drive up electricity congestion prices systematically, and in what way? Second, does that create a pricing pattern that market participants can exploit? This thesis uses seven years of hourly day-ahead and real-time locational marginal price data from PJM Interconnection, covering January 2019 through March 2026, along with operational data for 932 data center facilities, and a graph of the PJM transmission network. This thesis uses a difference-in-differences design, a network- based proximity measure, and a walk-forward trading model to answer both questions. The results show that data center load accumulation caused a significant and persistent congestion premium at grid nodes close by. It increased overall congestion prices, as well as the frequency of price spikes. This thesis creates theoretical models that estimate the congestion effect of any proposed data center based on its size and electrical proximity to a pricing node. While directional trading models were fairly accurate, adding in data center data did not substantially improve the models. The constant load and public nature of data centers makes it so that their information is reasonably priced in. These findings show that data center growth drastically changes price dynamics in affected areas. While no short-term trading edge is found from public DC data directly, the persistent congestion premium and elevated volatility off

  • Cross-Category Information Transmission and Price Comovement in Academy Awards Prediction Markets

    (2026-04-09) Gai, Emily; Holen, Margaret

    Prediction market research has largely focused on elections and other political events, leaving a gap for the exploration of other domains such as entertainment markets. This thesis investigates cross-category price transmission in the Academy Awards prediction markets on Polymarket, focusing on whether price changes in the Best Picture category spill over to contracts in other categories that are co-nominated under the same film. The analysis is conducted on daily prices across six major categories from 2025 and 2026. Using Principal Component Analysis, Granger causality, and pooled OLS regression, the study characterizes cross-category comovement, measures the lagged spillover effects, and decomposes the contract premium. The findings indicate that Best Picture price movements significantly predict changes in co-nominee prices within one day, with the effect amplified in more competitive seasons as co-nominee prices deviate from levels justified by their prior award records. This informs how cross-contract dependencies may affect pricing efficiency and contributes to the broader understanding of informational spillover dynamics in multi-category event markets.

  • Simulation-Based Optimization of MLB Pitching Rotations for Postseason Success

    (2026-04-09) Zhu, Nick; Akrotirianakis, Ioannis

    This thesis investigates how Major League Baseball (MLB) pitching rotations should be structured to maximize performance under the high-variance, short-series postseason format. Traditional roster construction methods focus on long-term regular season success, while existing postseason analyses treat games as independent events, ignoring sequencing effects. To address this gap, we develop a simulation-based framework that models outcomes at the game, series, and full playoff levels, and evaluate rotation strategies using Monte Carlo simulation. The resulting rotations are clustered into structural archetypes, revealing that top-heavy rotations achieve nearly equivalent performance at more cost-effective payrolls when compared to balanced rotations. Additionally, similar rotation costs across high-, medium-, and low-budget teams demonstrate diminishing returns to additional pitching investment. We then approximate championship probability with a surrogate model and optimize roster construction via integer programming, training the model on simulation outputs. The optimal solutions consistently feature young, surplus-value pitchers, highlighting the importance of internal development and timing in building postseason-optimized rotations.

  • Beyond Size: The Fragmentation of the Library Index and the Rise of Institutional Complexity in Academic Library Statistics

    (2026-04-09) Rillera, Tatiana Terese B.; Rigobon, Daniel

    For over four decades, academic libraries were quantitatively evaluated under the assumption that a single "Library Size" index could accurately summarize their operational capacity. This thesis mathematically tests that historical assumption by deploying longitudinal Principal Component Analysis and K-Means clustering on an eleven-year panel (2014-2015 to 2024-2025 Academic Year) of the federal IPEDS database. The analysis first uncovered the "IPEDS Imputation Trap" demonstrated that complex administrative reporting variables have artificially hijacked the primary variance of the national dataset due to federal imputation practices. After filtering this noise, the study proves that the historical library size monolith has permanently fractured. Projecting the national cohort onto these new axes establishes a modern taxonomy consisting of four distinct operational postures. Ultimately, this research proves that an academic library is no longer defined by a universal standard of size, but by the highly specialized economic survival strategies it executes within a decentralized landscape.

  • Numerical Simulations of Continuous-Time Flows for Monotone Games

    (2026-04-09) Ku, Charlie; Tangpi, Ludovic

    In this paper, we formulate three numerical algorithms for approximating Nash equilibria in monotone finite and mean-field games based on theoretical results that demonstrate convergence of ODE, Cesaro mean and gradient descent methods. Using a few example games, we compare convergence based on different game properties like strong monotonicity and symmetry; since our algorithms are based on Euler’s method, a linearization based on the Jacobian allows us to analyze the underlying structure of each game and how the error converges. In algorithms involving noise, we are also able to determine a more accurate noise floor based on this analysis.

  • Transit Networks & Housing Affordability: A Network-Based Econometric Analysis of Tract-Level Rent Burdens

    (2026-04-09) Lin, Erin; Rigobon, Daniel

    As nearly half of American renters spend more than 30 percent of their income on housing, cities have increasingly turned to public transit investment as a tool for expanding opportunity. The neighborhoods that gain the most connectivity, however, may be the ones that ultimately get priced out. This thesis investigates that tension empirically, using the 2019 Red Line Bus Rapid Transit expansion of Indianapolis's IndyGo system as a discrete network shock. Modeling the transit system as a graph and computing stop-level centrality measures before and after the expansion, this approach yields a tract-level measure of connectivity change that captures not just whether a neighborhood gained new stops, but how its position within the broader network shifted. This network-based approach, applied within a quasi-experimental framework, offers a more structurally grounded lens for studying transit's housing market consequences than conventional proximity-based methods allow. A Difference-in-Differences framework confirms that tracts experiencing the largest centrality gains exhibit persistently higher rent burden following the expansion, with effects that emerge roughly one year after the shock and remain stable over the subsequent four years. Notably, this pattern holds for centrality measures that capture global network position but not for degree centrality, pointing to a key mechanism: affordability pressure is driven not by the addition of nearby stops, but by a neighborhood becoming more deeply embedded in the broader transit network. Together, these findings suggest that the housing market consequences of transit investment are both broader and more structurally determined than conventional impact assessments recognize, with implications for how planners identify and respond to affordability pressure before it sets in.

  • When Does Sketching Suffice? Backward Stability in Randomized Least Squares

    (2026-04-09) Hsu, Ethan K.; Rebrova, Elizaveta

    Randomized sketching methods for overdetermined least-squares problems can reduce computational cost substantially relative to dense QR, but their numerical reliability is not yet fully understood in practice. This thesis studies when sketch-and-solve alone is sufficient to achieve backward stability and when iterative refinement is necessary. In particular, we compare sketch-and-precondition, sketch-and-precondition with iterative refinement (SPIR), and fast optimal stable sketchy iterative least squares (FOSSILS) across a systematic sweep of condition numbers κ ∈ {10^3 , . . . , 10^14}, aspect ratios m/n ∈ {10, . . . , 250}, residual sizes, and two noise models: one in which only b is perturbed and one in which both A and b are perturbed, for a total of 840 configurations. We complement the synthetic study with targeted follow-up experiments and real-data validation on SuiteSparse and LIBSVM matrices. The central empirical finding is that, within the tested regimes, the noise model is the main factor governing backward stability, while condition number and aspect ratio have little effect on pass rates. When only b is noisy, sketch-and-precondition with a warm start achieves backward stability in 96% of configurations, indicating that expensive iterative refinement is usually unnecessary. However, when A is also noisy, SPIR is the only method that remains reliably backward stable across the tested configurations. FOSSILS fails in about 10% of cases, and sketch-only methods fail throughout. These results provide practical guidance for selecting randomized least-squares solvers when both efficiency and certifiable numerical accuracy matter.

  • Finding the Right Group: Difficulty Measurement, False Group Analysis, and Automated Solving in NYT Connections

    (2026-04-09) Burda, Ben; Klusowski, Jason Matthew

    NYT Connections is a daily word-grouping puzzle that requires players to sort 16 words into four thematically linked groups of four. The puzzle’s difficulty arises from semantic ambiguity and deliberate misdirection, including “false groups” — sets of four words that form plausible but incorrect categories. This thesis analyzes Connections puzzles through three lenses using natural language processing methods. First, embedding-based metrics are developed to quantify puzzle difficulty. Group cohesion (mean within-group cosine similarity) and silhouette score correlate negatively with ground-truth difficulty ratings from the NYT, confirming that harder puzzles have less-separated semantic clusters. These correlations are modest, indicating that a substantial portion of perceived difficulty lies beyond what embeddings can capture. Second, false groups are formalized as a computable object in embedding space. A strong false group is defined as any four-word subset whose cohesion exceeds that of the weakest true group by a fixed margin δ. Applied to 226 labeled puzzles, this definition identifies strong false groups in 40.7% of puzzles. An overlap signature taxonomy reveals that near-miss groups — differing from a true solution by exactly one word — are the most common false group type at 46%. Third, an automated solver pipeline is developed and evaluated through a five-solver ablation study. Starting from a greedy baseline that achieves 0% solve rate, beam search, WordNet lexical augmentation, and iterative feedback are added incrementally. The full feedback-aware pipeline achieves a 15.2% full puzzle solve rate on a held-out test set of 46 puzzles, with top-1 accuracy of 56.5%.