Transparency
This project publishes its methodology, extraction prompts, sourcing policy and all registry data as open documents.
Published documents
| Document | Licence | Status |
|---|---|---|
| Collection methodology | CC-BY 4.0 | Working version |
| Extraction prompts | CC-BY 4.0 | Published |
| Sourcing policy | CC-BY 4.0 | Working version |
| Case record schema (JSON Schema) | CC-BY 4.0 | Published |
| Registry data (JSON) · CSV | CC-BY 4.0 | Updated with each release |
| Case-file texts and severity coding (JSON) | CC-BY-NC 4.0 | Updated with each release |
| Coding rules (summary of Codebook v0.5) | CC-BY 4.0 | Working version |
What is not published
The implementation code that runs the collection pipeline is privately maintained. The collection methodology and the extraction prompts describe every step it performs.
Use of AI in this catalogue
Language models are used to filter and structure, never to decide what is published. Articles are screened by Claude Haiku (Anthropic); candidate records are structured by Claude Sonnet (Anthropic); in October 2026 a batch of records was structured with Claude in an assisted review session. Every published record is checked by the author against its sources. Severity coding in the case-files is done by hand.
All transparency documents are licensed under CC-BY 4.0 unless otherwise noted.