Publication:

YASS: A Synthetic Dataset Curated for Large Language Model Generation of System Verilog Assertions via Distillation Methods

Loading...
Thumbnail Image

Files

Bolaji_Senior_Thesis_Spring_26.pdf (3.19 MB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Pre-silicon verification of System-on-Chip (SoC) designs consumes approximately 70% of overall development time, with SystemVerilog Assertions (SVA) serving as a critical tool for verifying design correctness. Manually authoring these assertions is time-consuming and error-prone, and recent efforts to automate SVA generation using Large Language Models (LLMs) have been limited by the scarcity of high-quality SVA training data. This thesis introduces YASS—named after its contributors’ initials—a synthetic SVA dataset created by using GPT-5 to generate assertions directly from 5,774 verified RTL designs in the VeriThoughts database without requiring natural language input at inference time. Generated assertions are filtered through a two-stage JasperGold pipeline consisting of syntax validation and formal verification, producing three quality-stratified dataset variants: all generated SVA, syntax-passing SVA, and formally verified assertions. Each variant is used to fine-tune Qwen2.5-Coder-7B-Instruct via QLoRA, producing three LoRA adapters evaluated against both the untrained Qwen2.5 baseline and GPT-5 on a held-out 578-design evaluation set. Results demonstrate that fine-tuning on YASS-derived data dramatically improves SVA generation quality: all adapters exceed 85% syntax pass rate compared to 37.2% for the untrained baseline. The Verified Assertions adapter achieves the highest proof rate (73.7%), surpassing GPT-5 (71.7%), while compiling the most designs (427 vs 335). A proposed composite quality score incorporating proof rate, counterexample penalty, coverage, and vacuity confirms these findings, with the Verified Assertions adapter achieving the highest A+B grade rate (63%) and the most Grade A outputs (194 designs, 45%). These results demonstrate that a 7B-parameter open-source model fine-tuned on quality-filtered data can match or exceed frontier model performance for SVA generation, and that dataset quality—not model size—is the primary driver of assertion generation accuracy. All code, data, and trained adapters are publicly available.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation