<mods:mods xmlns:mods="http://www.loc.gov/mods/v3" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" ID="etd984" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-2.xsd">
	<mods:titleInfo>
		<mods:title>Algorithms for Identifying Structural Variants in Human Genome</mods:title>
	</mods:titleInfo><mods:name type="personal">
		<mods:namePart>Ritz, Anna M</mods:namePart>
	<mods:role>
		<mods:roleTerm type="text">creator</mods:roleTerm>
	</mods:role>
	</mods:name>
<mods:originInfo>
	<mods:copyrightDate>2013</mods:copyrightDate>
</mods:originInfo>
<mods:physicalDescription>
        <mods:extent>17, 149 p.</mods:extent>
        <mods:digitalOrigin>born digital</mods:digitalOrigin>
</mods:physicalDescription>
<mods:note>Thesis (Ph.D. -- Brown University (2013)</mods:note>
<mods:name type="personal">
<mods:namePart>Raphael, Benjamin</mods:namePart>
<mods:role>
<mods:roleTerm type="text">Director</mods:roleTerm>
</mods:role>
</mods:name>

<mods:name type="personal">
<mods:namePart>Dunn, Casey</mods:namePart>
<mods:role>
<mods:roleTerm type="text">Reader</mods:roleTerm>
</mods:role>
</mods:name>

<mods:name type="personal">
<mods:namePart>Laidlaw, David</mods:namePart>
<mods:role>
<mods:roleTerm type="text">Reader</mods:roleTerm>
</mods:role>
</mods:name>

<mods:name type="personal">
<mods:namePart>Upfal, Eli</mods:namePart>
<mods:role>
<mods:roleTerm type="text">Reader</mods:roleTerm>
</mods:role>
</mods:name>
<mods:name type="corporate">
		<mods:namePart>Brown University. Computer Science</mods:namePart>
		<mods:role>
			<mods:roleTerm type="text">sponsor</mods:roleTerm>
		</mods:role>
		</mods:name>
	<mods:genre authority="aat">theses</mods:genre>
	<mods:subject>
        <mods:topic>structural variation</mods:topic>
    </mods:subject>

    <mods:subject>
        <mods:topic>array-comparative genomic hybridization</mods:topic>
    </mods:subject>

    <mods:subject>
        <mods:topic>DNA sequencing</mods:topic>
    </mods:subject>

	<mods:recordInfo>
		<mods:recordContentSource authority="marcorg">RPB</mods:recordContentSource>
		<mods:recordCreationDate encoding="iso8601">20131218</mods:recordCreationDate>        
	</mods:recordInfo>
<mods:language xmlns:xlink="http://www.w3.org/1999/xlink"><mods:languageTerm type="code" authority="iso639-2b">eng</mods:languageTerm><mods:languageTerm type="text">English</mods:languageTerm></mods:language><mods:abstract xmlns:xlink="http://www.w3.org/1999/xlink">Variation in genomes occurs in many forms, from single nucleotide changes to gains and losses of entire chromosomes. Large-scale rearrangements, called structural variants (SVs), are associated with numerous diseases and are common in cancer genomes. However, many SVs in mammalian genomes are found in highly repetitive regions, complicating their detection and characterization. Ongoing development of genomic technologies invite new algorithmic approaches to SV detection.&lt;br/&gt;&lt;br/&gt;

In this thesis, we present a collection of four algorithms that identify SVs using data from current and emerging genomic technologies. The first algorithm is designed for a technology called array-comparative genomic hybridization (aCGH), which measures the number of copies of DNA segments present in a test genome relative to a known reference genome. We describe a method to identify SVs that are common to a group of individuals, and apply the method to aCGH data from hundreds of cancer patients. We recover an SV in prostate cancer that is known to be  biologically important, and we infer a number of novel SVs in brain cancer.&lt;br/&gt;&lt;br/&gt;

Our other algorithms are designed for DNA sequencing technologies, which measure a broader range of SVs than aCGH data with the tradeoff of higher cost. One DNA sequencing technology, strobe sequencing, yields multiple sequences from a single, contiguous fragment of DNA. While strobes provide longer sequenced portions of DNA compared to other sequencing technologies, the per-base error rate is substantially higher. Our algorithms for SV detection exploit the benefits of multiply-linked DNA sequences while being robust to high sequencing error rates. We describe the first published method for SV detection using strobe sequencing, which finds the smallest number of SVs (relative to a known reference genome) that explain the strobes. We then improve upon our method with a probabilistic algorithm that better models the strobe sequencing data. Finally, we describe a de novo assembly algorithm for strobe sequencing data when a reference genome is unavailable. We assess the performance of these algorithms on simulated and real strobe sequencing data, and conclude that with appropriate algorithms, strobe sequencing compares favorably to other DNA sequencing technologies.</mods:abstract><mods:identifier xmlns:xlink="http://www.w3.org/1999/xlink" type="doi">10.7301/Z00863MD</mods:identifier><mods:accessCondition xmlns:xlink="http://www.w3.org/1999/xlink" type="rights statement" xlink:href="http://rightsstatements.org/vocab/InC/1.0/">In Copyright</mods:accessCondition><mods:accessCondition type="restriction on access">Collection is open for research.</mods:accessCondition><mods:typeOfResource xmlns:xlink="http://www.w3.org/1999/xlink" authority="primo">dissertations</mods:typeOfResource></mods:mods>