Skip to main content
뒤로

Human Gene Structure, Nomenclature, and Genomic Organization

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Genes and the Human Genome

Number of Genes in the Human Nuclear Genome

The human genome contains approximately 20,000 protein-coding genes, with the current official count at 19,901 according to the Wellcome-Sanger Institute. The definition of a gene is debated, but for clinical genomics, a gene is defined as a sequence of DNA nucleotides that codes for either a protein or a type of RNA, including its proximal upstream regulatory elements (promoter region).

  • Gene: Sequence of DNA coding for protein or RNA, including regulatory elements.

  • Protein-coding genes: Only about 1% of the human genome.

  • RNA-coding genes: Many types, not all code for proteins.

Gene Nomenclature

Gene naming is often inconsistent due to multiple discoverers and evolving knowledge about gene functions. The Human Genome Organization (HUGO) Gene Nomenclature Committee is the most authoritative body for gene names. Gene symbols are italicized, while protein symbols are not.

  • Gene symbol: Italicized (e.g., SHH for Sonic Hedgehog gene).

  • Protein symbol: Regular text (e.g., SHH protein).

  • Gene naming: Often based on function or discoverer; sometimes whimsical.

For more information, visit https://www.genenames.org/.

Gene Structure and Regulation

Structure of Protein-Coding Genes

Protein-coding genes have a characteristic structure, including regulatory regions, exons, and introns. Transcription factors and RNA polymerase II bind to the promoter region upstream of the transcriptional start site (TSS). The gene is transcribed into mRNA, which includes untranslated regions (UTRs) and coding regions (exons), while introns are spliced out.

  • Promoter region: Upstream regulatory sequence where transcription factors bind.

  • 5' UTR: Untranslated region at the start of mRNA.

  • Exons: Expressed sequences, coding for protein.

  • Introns: Non-coding sequences, spliced out during mRNA processing.

  • 3' UTR: Untranslated region at the end of mRNA.

  • Enhancers: Distant regulatory elements that promote transcription.

  • Insulators: Elements that block or suppress transcription.

Alternative splicing allows a single gene to produce multiple mRNA transcripts and protein products. The spliceosome is the complex responsible for splicing.

Example: The human genome has about 20,500 protein-coding genes but over 203,835 mRNA transcripts due to alternative splicing.

Schematic of a eukaryotic gene structure

Open Reading Frames (ORFs)

An Open Reading Frame (ORF) is a sequence of DNA that can be translated into a protein, starting with a start codon and ending with a stop codon. ORFs are important for identifying potential protein-coding regions in genomic sequences.

  • Start codon: Usually ATG in DNA.

  • Stop codon: TAA, TAG, or TGA in DNA.

Gene Representation and Genomic Visualization

Genomic Sequence Databases and Browsers

Genes are represented in databases at various levels of detail. Tools like the NCBI Map Viewer and UCSC Genome Browser allow visualization of gene locations, structures, and sequences. The ideogram provides a top-level view of chromosome structure and gene position.

  • Ideogram: Schematic representation of chromosomes, showing gene locations.

  • Gene length: Measured in base pairs (bp) or kilobases (Kb).

  • Genomic coordinates: Specify gene position on a chromosome.

Example: The EGFR gene is located at 7p11.2 on chromosome 7 and is 189,059 base pairs long.

EGFR gene location and structure on chromosome 7

Transcriptional Start Site and Nucleotide Sequence

Zooming in on a gene reveals the transcriptional start site (TSS) and the nucleotide sequence. The positive strand is depicted 5’ to 3’, and the negative strand is complementary. Genome browsers allow visualization of exons, introns, and raw DNA sequence.

  • Transcriptional start site (TSS): Where transcription begins.

  • Nucleotide sequence: DNA bases (A, T, C, G) shown in genomic viewers.

  • Strand orientation: Positive (5’ to 3’) and negative (3’ to 5’).

EGFR gene transcriptional start site and nucleotide sequence

Raw DNA Sequence Representation

Genomic sequences are often displayed in formats like GenBank "ORIGIN", showing base pair positions and nucleotide sequences for manual inspection or computational analysis.

  • Sequence format: Lowercase letters, grouped for readability.

  • Base pair numbering: Each line starts with a position number.

  • Strand: Usually positive strand is shown; negative strand can be deduced.

Raw DNA sequence in GenBank ORIGIN format

Gene Distribution and Organization

Gene Distribution in the Human Genome

The human genome is highly complex, with genes distributed unevenly. Some regions, called gene deserts, have few or no genes, while others contain clusters of related genes. Gene clusters often participate in similar biochemical pathways and may be co-regulated.

  • Gene deserts: Large regions with no genes.

  • Gene clusters: Groups of related genes, often co-regulated.

Overlapping and Nested Genes

Genes can overlap or be nested within each other, with transcription occurring in various directions. Overlapping genes may share regions, while nested genes are entirely contained within another gene’s sequence. These arrangements complicate gene regulation and annotation.

  • Overlapping genes: Two genes share part of their sequence.

  • Nested genes: One gene is located within another gene’s sequence.

  • Transcription direction: Genes can be transcribed in parallel or antiparallel directions.

Classification Table: Types of overlapping and nested genes based on overlap extent and transcription direction.

Type of Overlap

Direction of Transcription

Example

Partial

Convergent, Divergent, Parallel

Genes overlap partially, transcribed in same or opposite directions

Complete

Nested Antiparallel, Nested Parallel, Embedded Antiparallel, Embedded Parallel

One gene is fully contained within another, transcribed in same or opposite directions

Types of overlapping and nested genes

Summary

  • The human genome is composed of approximately 20,000 protein-coding genes, with complex regulatory and structural features.

  • Gene nomenclature and structure are essential for understanding genetic function and regulation.

  • Genomic visualization tools provide detailed views of gene location, structure, and sequence.

  • Genes can be distributed unevenly, with overlapping and nested arrangements adding complexity to genome architecture.

Pearson Logo

스터디 프렙