• DocumentCode
    2528807
  • Title

    ssahaSNP - a polymorphism detection tool on a whole genome scale

  • Author

    Ning, Zemin ; Caccamo, Mario ; Mullikin, James C.

  • Author_Institution
    Wellcome Trust Sanger Inst., Cambridge, UK
  • fYear
    2005
  • fDate
    8-11 Aug. 2005
  • Firstpage
    251
  • Lastpage
    252
  • Abstract
    We present a software package which can detect homozygous SNPs and indels on a eukaryotic genome scale from millions of shotgun reads. Matching seeds of a few kmer words are found to locate the position of the read on the genome. Full sequence alignment is performed to detect base variations. Quality values of both variation bases and neighbouring bases are checked to exclude possible sequence base errors. To analyze polymorphism level in the genome, we used the package to detect indels from 20 million WGS reads against the draft WGS assembly. From the dataset, we detected a total number of 663,660 indels, giving an estimated average indel density at about one indel every 2.48 kilobases. Distribution of indels length and variation of indel mapped times are also analyzed.
  • Keywords
    cellular biophysics; genetics; medical computing; molecular biophysics; software packages; base variation; draft whole genome shotgun assembly; eukaryotic genome scale; homozygous SNP; indel detection; kmer word; polymorphism detection tool; sequence alignment; shotgun read; single-nucleotide polymorphism; software package; ssahaSNP; Assembly; Bioinformatics; Cancer detection; Data structures; Databases; Diseases; Error analysis; Genomics; Packaging; Software packages;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computational Systems Bioinformatics Conference, 2005. Workshops and Poster Abstracts. IEEE
  • Print_ISBN
    0-7695-2442-7
  • Type

    conf

  • DOI
    10.1109/CSBW.2005.128
  • Filename
    1540618