Title Information
Title
Improving Information Propagation in Phylogenetic Workflows
Name: Personal
Name Part
Guang, August
Role
Role Term: Text
creator
Name: Personal
Name Part
Lawrence, Charles
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Dunn, Casey
Role
Role Term: Text
Reader
Name: Personal
Name Part
Lewis, Paul
Role
Role Term: Text
Reader
Name: Corporate
Name Part
Brown University. Department of Applied Mathematics
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2018
Physical Description
Extent
xiv, 100 p.
digitalOrigin
born digital
Note: thesis
Thesis (Ph. D.)--Brown University, 2018
Genre (aat)
theses
Abstract
Despite the enormous amount of biological variation and technical uncertainty in sequence data, most phylogenetic workflows propagate a single point estimate throughout the numerous analysis components, and only in a forward direction. This approach relies on three implicit assumptions: (i) the order of the analysis steps is biologically justified, (ii) a Markovian dependency structure exists between analysis components, and (iii) there is low relative entropy between results at each analysis step. There is evidence that these assumptions, in particular low relative entropy, are frequently violated in empirical studies with potential detrimental effects in phylogenetic analyses. In this thesis, I lay out a probabilistic framework that provides a unified perspective to provide context for evaluating priorities for future developments of methods and tools. I then develop a generative model of the natural and technical processes that produce observed genomic reads within the framework that can be used to assess and validate approaches that relax the implicit assumptions. Finally, I explore two ways to accommodate and propagate more information in a phylogenetic workflow. The first way, an HMM profile-sampling approach to genome assembly, relaxes the assumption of low relative entropy in results from the genome assembly analysis component. This approach finds relevant applications to HIV transmission networks. The second way, an iterative approach to identifying and resolving transcriptome assembly errors, capitalizes on the assumption of Markovian dependence.
Subject
Topic
HIV/AIDS
Subject
Topic
Probability
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/00871990")
Topic
Computational biology
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/01062326")
Topic
Phylogeny
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20180615
Identifier: DOI
10.26300/m4j5-dd88
Access Condition: rights statement (href="http://rightsstatements.org/vocab/InC/1.0/")
In Copyright
Access Condition: restriction on access
Collection is open for research.
Type of Resource (primo)
dissertations