Title Information
Title
Advancing the Semi-automation of Literature Identification for Evidence Synthesis: Evaluation Measure Prioritization, Training Data Development, and Language Model Implementation
Type of Resource (primo)
dissertations
Name: Personal
Name Part
Adam, Gaelen Phyfe
Role
Role Term: Text
creator
Name: Personal
Name Part
Balk, Ethan
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Trikalinos, Thomas
Role
Role Term: Text
Reader
Name: Personal
Name Part
Wallace, Byron
Role
Role Term: Text
Reader
Name: Personal
Name Part
Saldanha, Ian
Role
Role Term: Text
Reader
Name: Corporate
Name Part
Brown University. Department of Health Services, Policy and Practice
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2024
Physical Description
Extent
x, 73 p.
digitalOrigin
born digital
Note: thesis
Thesis (Ph. D.)--Brown University, 2024
Genre (aat)
theses
Abstract
Evidence synthesis products (e.g., systematic reviews, rapid reviews) form the basis of evidence-based healthcare. Literature identification—searching for studies and screening the resulting citations to identify relevant literature—is an important but time- and labor-intensive step in the systematic review process. Medical librarians serve an important role in developing search queries that balance the requirement to identify all relevant studies with the constraints of the team’s available time and budget, but this process is time-consuming, and not every review team has access to an experienced librarian. Thus, there is a clear need for computer-based tools that assist in the design and development of high-quality search queries. With the ongoing development of tools to semi-automate literature search query development, it is increasingly important to find meaningful ways to evaluate their performance. In the first of three sections of this dissertation, I describe a discrete choice experiment, conducted as a survey, to determine which types of measures a variety of stakeholders prefer. Surveyed systematic review methodologists and librarians would like to see studies report measures used to establish whether a tool identifies all relevant records as a ratio (i.e., sensitivity) and prefer measures that clearly indicate how much work it saves. In the second and third sections of this dissertation, I describe a project in which my colleagues and I divided a corpus of systematic review topics and search queries into training and validation sets. We subsequently developed an independent evaluation dataset. We trained language models on the training dataset and, after settling on all hyper-parameter settings using the validation dataset, evaluated them quantitatively, using the evaluation dataset. The models had a median sensitivity of 85% and required that the simulated team screen approximately 1,000 abstracts for every included citation. I also evaluated the models qualitatively, through semi-structured interviews with eight librarians, during which they piloted and evaluated the models on real search topics. The librarians generally expressed that although the tool-generated queries lacked both the necessary sensitivity and precision to be used without scrutiny, the queries could be used as teaching tools or as starting points for non-expert searchers.
Subject
Topic
Large Language Models
Subject
Topic
Systematic review,
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/00997916")
Topic
Library science
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20240501