Title Information
Title
Automated Text Mining to Improve the Curation of Genes Associated with Complex Disease
Name: Personal
Name Part
Superdock, Michael
Role
Role Term: Text
creator
Name: Personal
Name Part
Uzun, Alper
Role
Role Term: Text
creator
Name: Personal
Name Part
Sarkar, Indra Neil
Role
Role Term: Text
creator
Name: Personal
Name Part
Padbury, James
Role
Role Term: Text
creator
Type of Resource
text
Genre (aat)
posters
Origin Information
Date Created (keyDate="yes", encoding="w3cdtf")
2017
Language
Language Term: Code (ISO639-2B)
eng
Note (displayLabel="Scholarly concentration")
Biomedical Informatics
Note
All rights reserved
Abstract
Manual curation of primary literature is a common, time-intensive approach for identifying genes associated with a disease of interest. This project aims to minimize the workload of manual curation for genetic studies by semi-automating the curation process. A computational pipeline was created using text-mining techniques to extract genetic data and other distinguishing features from articles. Five predictive models were trained on these features to classify articles as "considered" or "not considered" for later review by curators. The models were evaluated against manual classifications of curated papers from the Database for Preeclampsia (dbPEC) and the Database for Preterm Birth (dbPTB). A Random Forest classifier performed best for both datasets, with an AUC of 0.825 for dbPEC articles and an AUC of 0.918 for dbPTB articles. This classifier had results consistent with a 32.5% workload reduction for the curation of dbPEC articles and a 79.6% workload reduction for the curation of dbPTB articles, while still capturing over 95% of validated genes.
Subject (LCSH)
Topic
Machine learning
Subject (LCSH)
Topic
Genetics
Subject (LCSH)
Topic
Premature infants