Transparency

This project publishes its methodology, extraction prompts, sourcing policy and all registry data as open documents.

Published documents

DocumentLicenceStatus
Collection methodologyCC-BY 4.0Working version
Extraction promptsCC-BY 4.0Published
Sourcing policyCC-BY 4.0Working version
Case record schema (JSON Schema)CC-BY 4.0Published
Registry data (JSON) · CSVCC-BY 4.0Updated with each release
Case-file texts and severity coding (JSON)CC-BY-NC 4.0Updated with each release
Coding rules (summary of Codebook v0.5)CC-BY 4.0Working version

What is not published

The implementation code that runs the collection pipeline is privately maintained. The collection methodology and the extraction prompts describe every step it performs.

Use of AI in this catalogue

Language models are used to filter and structure, never to decide what is published. Articles are screened by Claude Haiku (Anthropic); candidate records are structured by Claude Sonnet (Anthropic); in October 2026 a batch of records was structured with Claude in an assisted review session. Every published record is checked by the author against its sources. Severity coding in the case-files is done by hand.


All transparency documents are licensed under CC-BY 4.0 unless otherwise noted.