Publication:

Neurosymbolic Learning for Corporate Finance: A Cross-Domain Transfer Study

Loading...
Thumbnail Image

Files

Hidde_Lycklama_Thesis.pdf (300.9 KB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Dedhia et al. demonstrated that combining GraphMERT-based knowledge graph extraction with a bottom-up reasoning curriculum enables large language models to acquire domain-specific expertise in medicine. Whether this neurosymbolic pipeline generalizes to domains with structurally different ontologies remains an open question. Corporate finance, defined by deterministic financial identities and a pronounced terminological gap between formal ontology vocabulary and instructional language, provides an interesting test case. At a similarity threshold of 0.55, the FIBO seed ontology produces a fully degenerate initialization against the Berk & DeMarzo textbook, collapsing to only 18 usable triples, a failure mode absent in medicine, where ontology and corpus share a common technical register. Despite this, GraphMERT extracts 1,291 valid triples and we generate a 10,000-question curriculum to fine-tune QwQ-32B via LoRA. SFT yields asymmetric gains: +2.3pp on numerical reasoning and −4.7pp on conceptual reasoning, identifying curriculum composition as the primary engineering variable. GraphRAG evaluation confirms the knowledge graph statistically recovers the degradation caused by the degenerate seed (p = 0.039), while out-of-domain MMLU Finance results reveal an 18.4pp retrieval penalty, indicating textbook-derived KGs are more effective as training scaffolds than inference-time retrievers. Seed ontology alignment, not architectural incompatibility, is the principal barrier to turnkey deployment across new domains.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation