Lecture Notes in Genome Bioinformatics

Subhashini Srinivasan

The first and most comprehensive eight credit genome bioinformatics curriculum developed and delivered by Prof. Subhashini Srinivasan over the last 15 years at IBAB. The curriculum comprises 8 chapters presented as 22 chapter-wise podcasts, offering a complete journey through the principles, algorithms, tools, and applications of genome bioinformatics.

Episodes

  1. 6 days ago

    Chapter 4.1: Assembly Algorithms in Genome Bioinformatics

    Imagine a machine that could walk along the 2-meter-long DNA molecules packed inside the nucleus of every cell and report the identity of every base it encounters. There would be little need for genome assembly and much of this chapter would become unnecessary. Sequencing machines can accurately read only a limited stretch of DNA before the signal becomes too noisy. If a machine can reliably read 100–200 bases before it begins to “blabber,” the result is a short sequencing read, such as the approximately 150-base reads commonly generated by Illumina platforms. Now imagine deploying a billion Lilliputians, each landing at a random location on the nuclear DNA from many cells and walking as far as it can while accurately reporting the bases it encounters. We would obtain a billion short reads, each representing a small fragment of the genome. If these reads were distributed randomly across the genome, many would overlap with one another. These overlaps provide the clues needed to reconstruct progressively longer stretches of DNA, called contigs. This is the fundamental challenge of genome assembly: reconstructing a long DNA sequence from millions or billions of short, overlapping observations. The quality of an assembly is not simply an all-or-none measure. A perfect telomere-to-telomere (T2T) assembly is the ultimate goal of assembling genomes of any organism, but obtaining such an assembly can require substantially more data and sophisticated technologies. The human genome draft published in 2001 contained thousands of gaps, yet it was enormously valuable and transformed our ability to study genes, genomic variation, and the functional organization of the genome. Thus, a genome assembly does not need to be perfect to be useful. An assembly that produces sufficiently long contigs and scaffolds with meaningful genomic context can already provide a powerful foundation for gene discovery, comparative genomics, variant analysis, and aiding many biological applications.

About

The first and most comprehensive eight credit genome bioinformatics curriculum developed and delivered by Prof. Subhashini Srinivasan over the last 15 years at IBAB. The curriculum comprises 8 chapters presented as 22 chapter-wise podcasts, offering a complete journey through the principles, algorithms, tools, and applications of genome bioinformatics.