TY - JOUR AU - Karapetian, Karina AU - Jeon, Soo Min AU - Kwon, Jin-Won AU - Suh, Young-Kyoon PY - 2023 DA - 2023/3/8 TI - Supervised Relation Extraction Between Suicide-Related Entities and Drugs: Development and Usability Study of an Annotated PubMed Corpus JO - J Med Internet Res SP - e41100 VL - 25 KW - suicide KW - adverse drug events KW - information extraction KW - relation classification KW - bidirectional encoder representations from transformers KW - pharmacovigilance KW - natural language processing KW - PubMed KW - corpus KW - language model AB - Background: Drug-induced suicide has been debated as a crucial issue in both clinical and public health research. Published research articles contain valuable data on the drugs associated with suicidal adverse events. An automated process that extracts such information and rapidly detects drugs related to suicide risk is essential but has not been well established. Moreover, few data sets are available for training and validating classification models on drug-induced suicide. Objective: This study aimed to build a corpus of drug-suicide relations containing annotated entities for drugs, suicidal adverse events, and their relations. To confirm the effectiveness of the drug-suicide relation corpus, we evaluated the performance of a relation classification model using the corpus in conjunction with various embeddings. Methods: We collected the abstracts and titles of research articles associated with drugs and suicide from PubMed and manually annotated them along with their relations at the sentence level (adverse drug events, treatment, suicide means, or miscellaneous). To reduce the manual annotation effort, we preliminarily selected sentences with a pretrained zero-shot classifier or sentences containing only drug and suicide keywords. We trained a relation classification model using various Bidirectional Encoder Representations from Transformer embeddings with the proposed corpus. We then compared the performances of the model with different Bidirectional Encoder Representations from Transformer–based embeddings and selected the most suitable embedding for our corpus. Results: Our corpus comprised 11,894 sentences extracted from the titles and abstracts of the PubMed research articles. Each sentence was annotated with drug and suicide entities and the relationship between these 2 entities (adverse drug events, treatment, means, and miscellaneous). All of the tested relation classification models that were fine-tuned on the corpus accurately detected sentences of suicidal adverse events regardless of their pretrained type and data set properties. Conclusions: To our knowledge, this is the first and most extensive corpus of drug-suicide relations. SN - 1438-8871 UR - https://www.jmir.org/2023/1/e41100 UR - https://doi.org/10.2196/41100 UR - http://www.ncbi.nlm.nih.gov/pubmed/36884281 DO - 10.2196/41100 ID - info:doi/10.2196/41100 ER -