- Title Information
- Title
- Automated Text Mining to Improve the Curation of Genes Associated with Complex Disease
- Name:
Personal
- Name Part
- Superdock, Michael
- Role
- Role Term:
Text
- creator
- Name:
Personal
- Name Part
- Uzun, Alper
- Role
- Role Term:
Text
- creator
- Name:
Personal
- Name Part
- Sarkar, Indra Neil
- Role
- Role Term:
Text
- creator
- Name:
Personal
- Name Part
- Padbury, James
- Role
- Role Term:
Text
- creator
- Type of Resource
- text
- Genre (aat)
- posters
- Origin Information
- Date Created
(keyDate="yes", encoding="w3cdtf")
- 2017
- Language
- Language Term:
Code (ISO639-2B)
- eng
- Note
(displayLabel="Scholarly concentration")
- Biomedical Informatics
- Note
- All rights reserved
- Abstract
- Manual curation of primary literature is a common, time-intensive approach for identifying genes associated with a disease of interest. This project aims to minimize the workload of manual curation for genetic studies by semi-automating the curation process. A computational pipeline was created using text-mining techniques to extract genetic data and other distinguishing features from articles. Five predictive models were trained on these features to classify articles as "considered" or "not considered" for later review by curators. The models were evaluated against manual classifications of curated papers from the Database for Preeclampsia (dbPEC) and the Database for Preterm Birth (dbPTB). A Random Forest classifier performed best for both datasets, with an AUC of 0.825 for dbPEC articles and an AUC of 0.918 for dbPTB articles. This classifier had results consistent with a 32.5% workload reduction for the curation of dbPEC articles and a 79.6% workload reduction for the curation of dbPTB articles, while still capturing over 95% of validated genes.
- Subject (LCSH)
- Topic
- Machine learning
- Subject (LCSH)
- Topic
- Genetics
- Subject (LCSH)
- Topic
- Premature infants