Methodology & data strategy

Built on an analysis of over 300 federal investigation documents.

To create an empirical foundation for measuring safety vocabulary, we built a structured reference dataset from public federal catastrophe investigations.

The regulatory baseline

Starting with established safety culture frameworks.

We began with the Nuclear Regulatory Commission's safety culture attributes (IMC 0310 and NUREG-2165), the gold standard for categorizing human performance, decision making, and communication in high-reliability operations.

Corpus ingestion & document strategy

Parsing 300+ federal investigation reports.

We systematically ingested and analyzed full investigation reports across three major federal oversight bodies: the Chemical Safety Board (CSB), Nuclear Regulatory Commission (NRC), and National Transportation Safety Board (NTSB). Every extracted finding in our corpus is verified against official public records.

Mapping markers of civil disasters

Identifying recurring communication patterns.

By analyzing investigation findings across multi-employer job sites, utility grids, and transportation corridors, we isolated the specific language markers and organizational friction points that consistently precede civil disasters.

The natural language comparison engine

A structured reference dataset for natural language comparison.

This multi-agency corpus forms the ground-truth dataset for our diagnostic engine. When participants respond to the prompt in their own words, natural language processing compares their phrasing against this structured dataset, without relying on multiple-choice surveys. The result maps a crew's shared safety coverage, showing managers where they might spend time shoring up gaps.

Attribute definitions draw on public federal investigation and inspection literature, including CSB, NRC, and NTSB materials. Not affiliated with, reviewed by, or endorsed by any of these agencies.