Explore every episode of the podcast Lecture Notes in Genome Bioinformatics
| Title | Pub. Date | Duration | |
|---|---|---|---|
| Chapter 4.2: Gene Prediction | 14 Sep 2026 | 00:22:28 | |
For assembled genomes to become biologically useful resources, it is essential to identify and annotate the regions that encode proteins and other functional elements. It is important to note that only about 1.5% of the three-billion-base human genome consists of protein-coding sequences, while the majority comprises regulatory regions, introns, repetitive elements, non-coding RNAs, and other genomic features. In contrast, genomes of prokaryotes are much more compact, with a substantially higher proportion of coding DNA. Protein-coding genes in bacteria and archaea are often densely packed, with relatively short intergenic regions, and in some cases, genes may even overlap. Therefore, genome annotation, the process of identifying genes, predicting their structures, and assigning potential functions, is a critical step after assembly. While the quality of the assembly determines the accuracy with which genomic regions can be reconstructed, annotation transforms the raw sequence into a functional genome resource that can be used for comparative genomics, evolutionary studies, and understanding the genetic basis of biological traits.Β | |||
| Chapter 4.1: Assembly Algorithms in Genome Bioinformatics | 14 Sep 2026 | 00:22:34 | |
Imagine a machine that could walk along the 2-meter-long DNA molecules packed inside the nucleus of every cell and report the identity of every base it encounters. There would be little need for genome assembly and much of this chapter would become unnecessary. Sequencing machines can accurately read only a limited stretch of DNA before the signal becomes too noisy. If a machine can reliably read 100β200 bases before it begins to βblabber,β the result is a short sequencing read, such as the approximately 150-base reads commonly generated by Illumina platforms. Now imagine deploying a billion Lilliputians, each landing at a random location on the nuclear DNA from many cells and walking as far as it can while accurately reporting the bases it encounters. We would obtain a billion short reads, each representing a small fragment of the genome. If these reads were distributed randomly across the genome, many would overlap with one another. These overlaps provide the clues needed to reconstruct progressively longer stretches of DNA, called contigs. This is the fundamental challenge of genome assembly: reconstructing a long DNA sequence from millions or billions of short, overlapping observations. The quality of an assembly is not simply an all-or-none measure. A perfect telomere-to-telomere (T2T) assembly is the ultimate goal of assembling genomes of any organism, but obtaining such an assembly can require substantially more data and sophisticated technologies. The human genome draft published in 2001 contained thousands of gaps, yet it was enormously valuable and transformed our ability to study genes, genomic variation, and the functional organization of the genome. Thus, a genome assembly does not need to be perfect to be useful. An assembly that produces sufficiently long contigs and scaffolds with meaningful genomic context can already provide a powerful foundation for gene discovery, comparative genomics, variant analysis, and aiding many biological applications. | |||
| Chapter 8: Lecture Notes in Genome Bioinformatics | 12 Sep 2026 | 00:22:16 | |
This chapter chronicles the efforts of faculty members to build an experiential learning curriculum at IBAB when next-generation sequencing (NGS) was still in its infancy. Building research programs alongside teaching created an extraordinary opportunity for students to engage in real-world problem-solving at a time when computational tools were limited, yet scientific ambitions ran high. It describes the challenges of selecting and pursuing projects that addressed important biological and societal questionsβfrom tackling protein malnutrition and identifying the causative mutation underlying a rare familial disease to devising a strategy for controlling malaria. | |||
| Chapter 2 from Lecture Notes in Bioinformatics | 11 Sep 2026 | 00:21:29 | |
Understanding the genetic elements that govern biological function is critical for deciphering how perturbations in these elements can alter phenotype, leading to disease in humans or desirable traits in plants. Chapter 2 establishes the fundamental biological framework required to analyze high-throughput genomic data and link causative genotypes to biological phenotypes. | |||
| Chapter 7 from Lecture Notes in Genome Bioinformatics | 11 Sep 2026 | 00:23:47 | |
Demystifying the myth that modern AI will solve all biological problems and turn everyone into a bioinformatics-savvy biologist. This chapter exposes readers to the importance of understanding the concepts behind genome bioinformatics in the era of AI. | |||
| Chapter 1 from Lecture Notes in Genome Bioinformatics | 11 Sep 2026 | 00:20:36 | |
Introduction to the scope on the book. | |||