Q BankQuestion BankDocsDocuments

19.1.11—Sequence and protein databases

Syllabus
9700–2028–2029
Objective
19.1.11
Level
A2

Sequence and protein databases organise evidence

Bioinformatics organises, stores and analyses biological data such as genome sequences, expressed-gene data and protein amino-acid sequences. Database comparisons can reveal sequence similarity and help infer possible relatedness or model-organism relevance, but the result remains evidence rather than automatic proof of function.

  1. Define the question and check the input data: identify whether the query is a DNA sequence, expressed-gene dataset or protein sequence, and check that the sequence and annotation are sufficiently complete for the intended comparison.
  2. Search an appropriate reference database containing comparable sequence, expression or protein information. The database scope and quality determine what can be found and how relevant the matches are.
  3. Align or compare the query with reference entries using a suitable sequence-comparison algorithm. Record similarity and conserved regions rather than treating one short match as a complete interpretation.
  4. Apply the stated comparison settings and any supplied threshold or significance rule consistently; do not invent a universal cut-off. A result is conditional on the reference set, algorithm and data quality.
  5. Infer cautiously: conserved or highly similar sequence regions can support a possible common function or evolutionary relationship, but sequence homology does not by itself prove identical function, expression, phenotype or safety.
  6. Use independent biological or experimental evidence to test the proposed function or relationship before making a stronger claim.

Reference database — determines which known sequences or proteins can be compared.
Similarity/alignment — identifies matching or conserved regions and supports a hypothesis about relatedness or function.
Data quality and settings — incomplete, noisy or poorly matched inputs, together with algorithm or threshold choices, can change the apparent result.
Inference boundary — a database match is a starting point for interpretation, not a direct measurement of gene function or a replacement for validation.

Do not treat bioinformatics as the PCR, gel-electrophoresis or microarray laboratory workflow. No database name, software setting, numerical threshold or degree of similarity is assumed here; state only what the supplied data and reference comparison support.

ConceptA-Level CAIE Biology A2