19.1.11—Sequence and protein databases
- Syllabus
- 9700–2028–2029
- Objective
- 19.1.11
- Level
- A2
Bioinformatics organises, stores and analyses biological data such as genome sequences, expressed-gene data and protein amino-acid sequences. Database comparisons can reveal sequence similarity and help infer possible relatedness or model-organism relevance, but the result remains evidence rather than automatic proof of function.
Reference database — determines which known sequences or proteins can be compared.
Similarity/alignment — identifies matching or conserved regions and supports a hypothesis about relatedness or function.
Data quality and settings — incomplete, noisy or poorly matched inputs, together with algorithm or threshold choices, can change the apparent result.
Inference boundary — a database match is a starting point for interpretation, not a direct measurement of gene function or a replacement for validation.
Do not treat bioinformatics as the PCR, gel-electrophoresis or microarray laboratory workflow. No database name, software setting, numerical threshold or degree of similarity is assumed here; state only what the supplied data and reference comparison support.