Description
This book provides a comprehensive treatment of the principles, algorithms, and standards that underpin genomic data compression and processing. The first chapters introduce the fundamental roles of compression, key metrics such as compression ratio and runtime, and criteria for selecting appropriate tools. The book then outlines the main genomic processing pipelines, with a description of the data types arising at each stage. The compression of the different data types—FASTA, FASTQ, SAM/BAM, VCF, and single-cell transcriptomic data—are examined in detail in the following chapters, highlighting both algorithmic innovations and empirical benchmarks. Standardization efforts are also discussed, focusing on the ISO/IEC 23092 (MPEG-G) series and the Global Alliance for Genomics and Health (GA4GH). One of the last chapters of the book explores applications of compression beyond storage, with particular focus on the computational advantages of using lossless compressors for classification of genomic data. The book concludes with learning-based compression methods, highlighting deep neural architectures, model mixing, and autoencoder frameworks representing the current state of the art in genomic data representation.
