<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://teaching.healthtech.dtu.dk/22111/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Henni</id>
	<title>22111 - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://teaching.healthtech.dtu.dk/22111/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Henni"/>
	<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php/Special:Contributions/Henni"/>
	<updated>2026-09-08T10:26:53Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=999</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=999"/>
		<updated>2026-09-07T08:48:11Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the button labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. &amp;lt;!-- It looks like this: [[File:Download.png]]. --&amp;gt;Click it, and then set &amp;lt;u&amp;gt;Dataset&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Format&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and click the new &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; button. This will open the text format in a new tab. What you see here basically contains all the information you have seen in the graphical interface (except for some information that is retrieved from cross-links, such as the 3D structure).&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try to answer the following questions. Try each question twice, &#039;&#039;first&#039;&#039; by thinking about the question and using the Advanced Search interface and/or editing the search string, &#039;&#039;second&#039;&#039; by prompting your favorite chatbot (Gemini, ChatGPT, Copilot, Claude, etc.). If the answers differ wildly, try some prompt engineering to make them converge.&lt;br /&gt;
&lt;br /&gt;
Regarding the first approach, remember to write your search string in the answer. Regarding the second approach, write the exact prompts you used and state which AI you worked with.&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name and/or family into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=998</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=998"/>
		<updated>2026-09-06T16:21:19Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&amp;lt;br&amp;gt;&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
* ChatGPT (free version): 9,080 (wrong). Asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the wrong answer &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: This is a question that does not have one single correct answer. A good attempt could be:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Using the prompt &amp;quot;How would you find insulin from as many organisms as possible from UniProt, without including entries that are not insulin?&amp;quot;:&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;family:&amp;quot;insulin family&amp;quot; AND gene:ins AND length:[50 TO 150]&amp;lt;/tt&amp;gt; which gives 661 hits (Not bad).&lt;br /&gt;
* Copilot: &lt;br /&gt;
 ((protein_name:&amp;quot;Insulin&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-1&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-2&amp;quot;))&lt;br /&gt;
 NOT (&lt;br /&gt;
    protein_name:&amp;quot;Insulin receptor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-like growth factor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-induced&amp;quot;&lt;br /&gt;
 )&lt;br /&gt;
:which gives 10,353 hits with way too many false positives. But it also suggests &amp;lt;tt&amp;gt;family:&amp;quot;Insulin family&amp;quot;&amp;lt;/tt&amp;gt; (8,929 hits) and combining that with filtering on protein name. &lt;br /&gt;
* ChatGPT (free version): does not come up with a useful UniProt query string, but suggests a lengthy workflow involving BLAST.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &lt;br /&gt;
* Our answer: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 15 hits.&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;(gene:HBA* OR family:&amp;quot;hemoglobin alpha-type family&amp;quot;) AND taxonomy_id:9895&amp;lt;/tt&amp;gt;, 26 hits, actually better than our answer, because some entries were named &amp;quot;Hemoglobin subunit alpha&amp;quot; &#039;&#039;without&#039;&#039; &amp;quot;alpha globin&amp;quot; as an alternative name.&lt;br /&gt;
* Copilot: &lt;br /&gt;
 taxonomy_id:9845 AND&lt;br /&gt;
 (protein_name:&amp;quot;Hemoglobin subunit alpha&amp;quot; OR gene_exact:hba1 OR gene_exact:hba2 OR gene_exact:hba)&lt;br /&gt;
:giving 37 hits without false positives, even better.&lt;br /&gt;
* ChatGPT (free version): required two prompts before it came up with the query string &amp;lt;tt&amp;gt;(protein_name:&amp;quot;hemoglobin subunit alpha&amp;quot; OR protein_name:&amp;quot;alpha-globin&amp;quot;) AND taxonomy_id:9845&amp;lt;/tt&amp;gt;, giving 58 hits including the &amp;quot;Alpha-globin transcription factor CP2&amp;quot; false positives.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;gene:HBA* AND organism:&amp;quot;Columba livia&amp;quot;&amp;lt;/tt&amp;gt; which does not work. However, it also suggests replacing the last part with &amp;lt;tt&amp;gt;taxonomy_id:8932&amp;lt;/tt&amp;gt;, and then you get 3 hits. Whether the last TrEMBL hit (&amp;quot;Hemoglobin subunit alpha-D-like&amp;quot;) is a false positive can be discussed.&lt;br /&gt;
* Copilot: gives several suggestion, of which only some are functional, e.g. &amp;lt;tt&amp;gt;organism_id:8932 AND (gene:hbaa OR gene:hbad OR protein_name:&amp;quot;alpha-A globin&amp;quot; OR protein_name:&amp;quot;alpha-D globin&amp;quot;)&amp;lt;/tt&amp;gt; (3 hits).&lt;br /&gt;
* ChatGPT (free version): &amp;lt;tt&amp;gt;organism_name:&amp;quot;Columba livia&amp;quot; AND (gene_exact:HBAA OR gene_exact:HBAD)&amp;lt;/tt&amp;gt; (3 hits).&lt;br /&gt;
Note that you can shorten the last suggestion to &amp;lt;tt&amp;gt;organism_name:&amp;quot;Columba livia&amp;quot; AND (gene_exact:HBA*)&amp;lt;/tt&amp;gt;&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=997</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=997"/>
		<updated>2026-09-06T15:59:37Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&amp;lt;br&amp;gt;&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
* ChatGPT (free version): 9,080 (wrong). Asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the wrong answer &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: This is a question that does not have one single correct answer. A good attempt could be:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Using the prompt &amp;quot;How would you find insulin from as many organisms as possible from UniProt, without including entries that are not insulin?&amp;quot;:&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;family:&amp;quot;insulin family&amp;quot; AND gene:ins AND length:[50 TO 150]&amp;lt;/tt&amp;gt; which gives 661 hits (Not bad).&lt;br /&gt;
* Copilot: &lt;br /&gt;
 ((protein_name:&amp;quot;Insulin&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-1&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-2&amp;quot;))&lt;br /&gt;
 NOT (&lt;br /&gt;
    protein_name:&amp;quot;Insulin receptor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-like growth factor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-induced&amp;quot;&lt;br /&gt;
 )&lt;br /&gt;
:which gives 10,353 hits with way too many false positives. But it also suggests &amp;lt;tt&amp;gt;family:&amp;quot;Insulin family&amp;quot;&amp;lt;/tt&amp;gt; (8,929 hits) and combining that with filtering on protein name. &lt;br /&gt;
* ChatGPT (free version): does not come up with a useful UniProt query string, but suggests a lengthy workflow involving BLAST.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &lt;br /&gt;
* Our answer: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 15 hits.&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;(gene:HBA* OR family:&amp;quot;hemoglobin alpha-type family&amp;quot;) AND taxonomy_id:9895&amp;lt;/tt&amp;gt;, 26 hits, actually better than our answer, because some entries were named &amp;quot;Hemoglobin subunit alpha&amp;quot; &#039;&#039;without&#039;&#039; &amp;quot;alpha globin&amp;quot; as an alternative name.&lt;br /&gt;
* Copilot: &lt;br /&gt;
 taxonomy_id:9845 AND&lt;br /&gt;
 (protein_name:&amp;quot;Hemoglobin subunit alpha&amp;quot; OR gene_exact:hba1 OR gene_exact:hba2 OR gene_exact:hba)&lt;br /&gt;
:giving 37 hits without false positives, even better.&lt;br /&gt;
* ChatGPT (free version): required two prompts before it came up with the query string &amp;lt;tt&amp;gt;(protein_name:&amp;quot;hemoglobin subunit alpha&amp;quot; OR protein_name:&amp;quot;alpha-globin&amp;quot;) AND taxonomy_id:9845&amp;lt;/tt&amp;gt;, giving 57 hits inlcuding the &amp;quot;Alpha-globin transcription factor CP2&amp;quot; false positives.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=996</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=996"/>
		<updated>2026-09-06T15:20:55Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the button labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. &amp;lt;!-- It looks like this: [[File:Download.png]]. --&amp;gt;Click it, and then set &amp;lt;u&amp;gt;Dataset&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Format&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and click the new &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; button. This will open the text format in a new tab. What you see here basically contains all the information you have seen in the graphical interface (except for some information that is retrieved from cross-links, such as the 3D structure).&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try to answer the following questions. Try each question twice, &#039;&#039;first&#039;&#039; by thinking about the question and using the Advanced Search interface and/or editing the search string, &#039;&#039;second&#039;&#039; by prompting your favorite chatbot (Gemini, ChatGPT, Copilot, Claude, etc.). If the answers differ wildly, try some prompt engineering to make them converge.&lt;br /&gt;
&lt;br /&gt;
Regarding the first approach, remember to write your search string in the answer. Regarding the second approach, write the exact prompts you used.&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name and/or family into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=995</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=995"/>
		<updated>2026-09-04T15:32:02Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&amp;lt;br&amp;gt;&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
* ChatGPT (free version): 9,080 (wrong). Asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the wrong answer &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: This is a question that does not have one single correct answer. A good attempt could be:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Using the prompt &amp;quot;How would you find insulin from as many organisms as possible from UniProt, without including entries that are not insulin?&amp;quot;:&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;family:&amp;quot;insulin family&amp;quot; AND gene:ins AND length:[50 TO 150]&amp;lt;/tt&amp;gt; which gives 661 hits. Not bad.&lt;br /&gt;
* Copilot: &lt;br /&gt;
 ((protein_name:&amp;quot;Insulin&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-1&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-2&amp;quot;))&lt;br /&gt;
 NOT (&lt;br /&gt;
    protein_name:&amp;quot;Insulin receptor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-like growth factor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-induced&amp;quot;&lt;br /&gt;
 )&lt;br /&gt;
which gives 10,353 hits with way too many false positives. But it also suggests &amp;lt;tt&amp;gt;family:&amp;quot;Insulin family&amp;quot;&amp;lt;/tt&amp;gt; (8,929 hits) and combining that with filtering on protein name. &lt;br /&gt;
* ChatGPT (free version): does not come up with a useful UniProt query string, but suggests a lengthy workflow involving BLAST.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 15 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=994</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=994"/>
		<updated>2026-09-04T15:29:15Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&amp;lt;br&amp;gt;&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
* ChatGPT (free version): 9,080 (wrong). Asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the wrong answer &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: This is a question that does not have one single correct answer. A good attempt could be:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Using the prompt &amp;quot;How would you find insulin from as many organisms as possible from UniProt, without including entries that are not insulin?&amp;quot;:&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;family:&amp;quot;insulin family&amp;quot; AND gene:ins AND length:[50 TO 150]&amp;lt;/tt&amp;gt; which gives 661 hits. Not bad.&lt;br /&gt;
* Copilot: &lt;br /&gt;
 ((protein_name:&amp;quot;Insulin&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-1&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-2&amp;quot;))&lt;br /&gt;
 NOT (&lt;br /&gt;
    protein_name:&amp;quot;Insulin receptor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-like growth factor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-induced&amp;quot;&lt;br /&gt;
 )&lt;br /&gt;
which gives 10,353 hits with way too many false positives. But it also suggests &amp;lt;tt&amp;gt;family:&amp;quot;Insulin family&amp;quot;&amp;lt;/tt&amp;gt; (8,929 hits) and combining that with filtering on protein name. &lt;br /&gt;
* ChatGPT (free version): does not come up with a useful UniProt query string, but suggests a lengthy workflow involving BLAST.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=993</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=993"/>
		<updated>2026-09-04T15:28:02Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&amp;lt;br&amp;gt;&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
* ChatGPT (free version): 9,080 (wrong). Asking it to provide the &amp;quot;UniProt query syntax&amp;quot; gave the wrong answer &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: This is a question that does not have one single correct answer. A good attempt could be:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Using the prompt &amp;quot;How would you find insulin from as many organisms as possible from UniProt, without including entries that are not insulin?&amp;quot;:&lt;br /&gt;
* Google: &amp;lt;tt&amp;gt;family:&amp;quot;insulin family&amp;quot; AND gene:ins AND length:[50 TO 150]&amp;lt;/tt&amp;gt; which gives 661 hits. Not bad.&lt;br /&gt;
* Copilot: &lt;br /&gt;
 ((protein_name:&amp;quot;Insulin&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-1&amp;quot;) OR&lt;br /&gt;
 (protein_name:&amp;quot;Insulin-2&amp;quot;))&lt;br /&gt;
 NOT (&lt;br /&gt;
    protein_name:&amp;quot;Insulin receptor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-like growth factor&amp;quot; OR&lt;br /&gt;
    protein_name:&amp;quot;Insulin-induced&amp;quot;&lt;br /&gt;
 )&lt;br /&gt;
which gives 10,353 hits with way too many false positives. But it also suggests &amp;lt;tt&amp;gt;family:&amp;quot;Insulin family&amp;quot;&amp;lt;/tt&amp;gt; (8,929 hits). &lt;br /&gt;
* ChatGPT (free version): does not come up with a useful UniProt query string, but suggests a lengthy workflow involving BLAST.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=992</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=992"/>
		<updated>2026-09-04T14:13:09Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the&amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries &amp;lt;tt&amp;gt;organism_id:83334&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;taxonomy_id:83334&amp;lt;/tt&amp;gt; are equivalent. They are not (try for yourself).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=991</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=991"/>
		<updated>2026-09-04T14:12:07Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own, with and without AI */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot; (Fair enough).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;The raw total climbs into the &#039;&#039;&#039;hundreds of thousands of protein entries&#039;&#039;&#039;, primarily sitting unreviewed within TrEMBL&amp;quot; (!). However, asking it to provide the&amp;quot;UniProt query syntax&amp;quot; gave the correct search string.&lt;br /&gt;
* Copilot: 5,999 (correct). However, it falsely claims that the queries organism_id:83334 and taxonomy_id:83334 are equivalent. They are not (try for yourself).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=990</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=990"/>
		<updated>2026-09-04T13:53:24Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.1&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 68,346 hits.&lt;br /&gt;
&lt;br /&gt;
Using the question verbatim as a prompt yielded:&lt;br /&gt;
* Google: &amp;quot;Searching broadly for all &#039;&#039;E. coli&#039;&#039; strains captures a pool of &#039;&#039;&#039;30 to 45 million entries&#039;&#039;&#039;&amp;quot; (!).&lt;br /&gt;
* Copilot: 68,344 UniProtKB protein entries (that&#039;s an error of 2).&lt;br /&gt;
* ChatGPT (free version): &amp;quot;The exact query to run in UniProt is: &amp;lt;tt&amp;gt;taxonomy_id:562&amp;lt;/tt&amp;gt;&amp;quot;. Fair enough.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.2&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 5,999 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.3&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.4&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4.5&#039;&#039;&#039;: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=989</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=989"/>
		<updated>2026-09-04T13:32:49Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* On your own */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the button labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. &amp;lt;!-- It looks like this: [[File:Download.png]]. --&amp;gt;Click it, and then set &amp;lt;u&amp;gt;Dataset&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Format&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and click the new &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; button. This will open the text format in a new tab. What you see here basically contains all the information you have seen in the graphical interface (except for some information that is retrieved from cross-links, such as the 3D structure).&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own, with and without AI ==&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try to answer the following questions. Try each question twice, &#039;&#039;first&#039;&#039; by thinking about the question and using the Advanced Search interface and/or editing the search string, &#039;&#039;second&#039;&#039; by prompting your favorite chatbot (Gemini, ChatGPT, Copilot, Claude, etc.). If the answers differ wildly, try some prompt engineering to make them converge.&lt;br /&gt;
&lt;br /&gt;
Regarding the first approach, remember to write your search string in the answer. Regarding the second approach, write the exact prompts you used.&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=988</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=988"/>
		<updated>2026-09-04T10:26:41Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Text format */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the button labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. &amp;lt;!-- It looks like this: [[File:Download.png]]. --&amp;gt;Click it, and then set &amp;lt;u&amp;gt;Dataset&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Format&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and click the new &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; button. This will open the text format in a new tab. What you see here basically contains all the information you have seen in the graphical interface (except for some information that is retrieved from cross-links, such as the 3D structure).&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=987</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=987"/>
		<updated>2026-09-03T16:11:44Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Text format */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.7:&#039;&#039;&#039; Find the two feature table lines that describe the signal peptide and copy them to your answer.&lt;br /&gt;
 FT   SIGNAL          1..24&lt;br /&gt;
 FT                   /evidence=&amp;quot;ECO:0000269|PubMed:14426955&amp;quot;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=986</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=986"/>
		<updated>2026-09-03T16:04:12Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Text format */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the button labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. &amp;lt;!-- It looks like this: [[File:Download.png]]. --&amp;gt;Click it, and then set &amp;lt;u&amp;gt;Dataset&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Format&amp;lt;/u&amp;gt; to &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and click the new &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; button. This will open the text format in a new tab. What you see here basically contains all the information you have seen in the graphical interface (except for some information that is retrieved from cross-links, such as the 3D structure).&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=985</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=985"/>
		<updated>2026-09-03T15:52:50Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* The contents of UniProt */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:How many mRNA sequences can you find links to, and what are their identifiers?&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the cross-links to PDB, and you can select which of the cross-linked structures to show by clicking on one of the lines. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. Unlike the nucleotide databases, however, UniProt only provides direct links to one of them, namely PDBe. &amp;lt;!-- We will work with the American version of PDB later in the course. --&amp;gt;As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined many times (under different experimental conditions). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039;&lt;br /&gt;
:As a default, UniProt displays the first structure in the list. This one shows several copies of the A and B chains. Try clicking the subsequent rows (&#039;&#039;&#039;Note:&#039;&#039;&#039; don&#039;t click the PDB identifier, just click somewhere in the row) until you find a structure that only shows one A chain and one B chain. What is the PDB Identifier of that structure? Insert a screenshot of the 3D structure in your answer.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039;&lt;br /&gt;
What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=984</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=984"/>
		<updated>2026-09-03T15:40:31Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Other databases linked from UniProt */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else? &amp;lt;br&amp;gt;IPR004825.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:PDB9R4E.png&amp;diff=983</id>
		<title>File:PDB9R4E.png</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:PDB9R4E.png&amp;diff=983"/>
		<updated>2026-09-03T15:38:39Z</updated>

		<summary type="html">&lt;p&gt;Henni: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=982</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=982"/>
		<updated>2026-09-03T15:38:06Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* The contents of UniProt */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; How many mRNA sequences can you find links to, and what are their identifiers? &amp;lt;br&amp;gt;There are four mRNA sequences linked, and they are called: X70508, AY899304, BT006808, and BC005255.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.5:&#039;&#039;&#039; What is the PDB Identifier of that structure? &amp;lt;br&amp;gt;The first structure to contain only one A and one B chain is &#039;&#039;&#039;9R4E&#039;&#039;&#039;, number 3 in the list.&lt;br /&gt;
Screenshot is here:&lt;br /&gt;
&lt;br /&gt;
[[Image:PDB9R4E.png]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.6:&#039;&#039;&#039; What is the InterPro identifier of the family that is just called &amp;quot;Insulin&amp;quot; and nothing else?&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=981</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=981"/>
		<updated>2026-09-02T14:51:55Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Advanced search */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;12,497,853, of these 45,490 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,922, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;742&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;903 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;5,341, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;9,009 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,324 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;889 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=980</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=980"/>
		<updated>2026-09-02T14:43:22Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* About strains and subspecies */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the actual cross-links to PDB. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. We will work with the American version of PDB later in the course. As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined several times. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of hits may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;or expand search to &amp;quot;1423&amp;quot; to &amp;lt;u&amp;gt;include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=979</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=979"/>
		<updated>2026-09-02T14:28:13Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Advanced search */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the actual cross-links to PDB. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. We will work with the American version of PDB later in the course. As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined several times. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Experimental&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of entries in Swiss-Prot may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;&amp;lt;u&amp;gt;or expand search to &amp;quot;1423&amp;quot; to include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=978</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=978"/>
		<updated>2026-09-02T14:18:34Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* The contents of UniProt */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;37&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;18,656,301, of these 45,100 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,891, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;734&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;18,297 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;35,762, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;46,068 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,350 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;905 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=977</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=977"/>
		<updated>2026-09-02T13:59:53Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Simple text mining */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just mentioning the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Popular organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advanced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the actual cross-links to PDB. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. We will work with the American version of PDB later in the course. As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined several times. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Any experimental assertion&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of entries in Swiss-Prot may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;&amp;lt;u&amp;gt;or expand search to &amp;quot;1423&amp;quot; to include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=976</id>
		<title>ExUniProt-answers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExUniProt-answers&amp;diff=976"/>
		<updated>2026-09-02T13:59:39Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Simple text mining */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
The numbers are found using UniProt 2025_03 on Sep 16, 2025&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
# How many hits do you find? &amp;lt;br&amp;gt;35,870&lt;br /&gt;
# How many of these hits are from Swiss-Prot? &amp;lt;br&amp;gt;1,774&lt;br /&gt;
# Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? &amp;lt;br&amp;gt;It&#039;s P01308 / INS_HUMAN (not necessarily the top hit, but still on the first page).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;2,375 and 1,157&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;217 and 60, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039; How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;68 and 25, search string: &amp;lt;tt&amp;gt;(organism_id:9606) AND (protein_name:insulin) NOT (protein_name:insulin-like)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039;&lt;br /&gt;
# How did you do this? &amp;lt;br&amp;gt;by adding &amp;lt;tt&amp;gt;NOT (protein_name:receptor)&amp;lt;/tt&amp;gt; to the query box.&lt;br /&gt;
# How many hits are now left? How many of these are from Swiss-Prot? &amp;lt;br&amp;gt;49 and 16&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039;&lt;br /&gt;
# How many references are there in the insulin entry? &amp;lt;br&amp;gt;36&lt;br /&gt;
# Why do you think insulin is such a highly investigated protein? &amp;lt;br&amp;gt;Because it is linked to a common and serious disease (diabetes) and used as a drug.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
# Where do you find insulin? &amp;lt;br&amp;gt;It is secreted from the cell (this is written just below the section heading. Under  &amp;lt;u&amp;gt;GO - Cellular component&amp;lt;/u&amp;gt; you can find additional locations mentioned, such as &amp;lt;u&amp;gt;endoplasmic reticulum lumen&amp;lt;/u&amp;gt;, but these are temporary stages on the way to secretion).&lt;br /&gt;
# Why do you think is it found there? &amp;lt;br&amp;gt;Because it is a hormone - it has to travel through the bloodstream to influence other cells.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039; How long is the signal peptide and the propeptide, respectively? &amp;lt;br&amp;gt;24 and 31 amino acids.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039; Which positions are in β-sheet conformation in insulin?  &amp;lt;br&amp;gt;Positions 26-29, 48-50, 56-58, 74-76, and 98-101.&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;No questions asked here.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; How many proteins did you find, and what was the search string (the text in the search field)? &amp;lt;br&amp;gt;18,656,301, of these 45,100 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal:*)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; How many proteins do you find now, and what has the search string changed into? &amp;lt;br&amp;gt;3,891, they are &#039;&#039;all&#039;&#039; from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*)&amp;lt;/tt&amp;gt;&amp;lt;br&amp;gt;Note that the &amp;quot;experimental&amp;quot; evidence is only found in Swiss-Prot entries, not in TrEMBL!&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; How many proteins do you find now, and what is the search string? &amp;lt;br&amp;gt;734&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(ft_signal_exp:*) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? &amp;lt;br&amp;gt;18,297 results, of these only 63 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(organism_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? &amp;lt;br&amp;gt;35,762, of these 4,280 from Swiss-Prot&amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(taxonomy_id:1423)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; How many proteins of maximum length 10 do you find? &amp;lt;br&amp;gt;46,068 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10])&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;1,350 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; How many proteins are now left? &amp;lt;br&amp;gt;905 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; How many human non-fragment proteins of maximum length 10 do you find in UniProt? &amp;lt;br&amp;gt;6 &amp;lt;br&amp;gt;&amp;lt;tt&amp;gt;(length:[1 TO 10]) AND (existence:1) AND (fragment:false) AND (organism_id:9606)&amp;lt;/tt&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Here they are in FASTA format:&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;sp|P0DPR3|TRDD1_HUMAN T cell receptor delta diversity 1 OS=Homo sapiens OX=9606 GN=TRDD1 PE=1 SV=1&lt;br /&gt;
 EI&lt;br /&gt;
 &amp;gt;sp|P01858|TUFT_HUMAN Phagocytosis-stimulating peptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 TKPR&lt;br /&gt;
 &amp;gt;sp|P02729|GLUR_HUMAN Urine glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEHSHDGA&lt;br /&gt;
 &amp;gt;sp|P01358|GAJU_HUMAN Gastric juice peptide 1 OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 LAAGKVEDSD&lt;br /&gt;
 &amp;gt;sp|P02728|GLEM_HUMAN Erythrocyte membrane glycopeptide OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 CEGHSHDHGA&lt;br /&gt;
 &amp;gt;sp|P22103|PNEU_HUMAN Pneumadin OS=Homo sapiens OX=9606 PE=1 SV=1&lt;br /&gt;
 AGEPKLDAGV&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.1: &amp;lt;tt&amp;gt;(taxonomy_id:562)&amp;lt;/tt&amp;gt;, 762,443 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.2: &amp;lt;tt&amp;gt;(taxonomy_id:83334)&amp;lt;/tt&amp;gt;, 18,197 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.3: &amp;lt;tt&amp;gt;(protein_name:insulin) AND (gene:ins) NOT (protein_name:insulin-like) NOT (protein_name:&amp;quot;insulin related&amp;quot;)&amp;lt;/tt&amp;gt;, 602 hits&amp;lt;br&amp;gt;&lt;br /&gt;
(This is a question that does not have one single correct answer)&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.4: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha globin&amp;quot;) AND (taxonomy_id:9845) NOT (protein_name:&amp;quot;transcription factor&amp;quot;)&amp;lt;/tt&amp;gt;, 17 hits.&lt;br /&gt;
&lt;br /&gt;
QUESTION 4.5: &amp;lt;tt&amp;gt;(protein_name:&amp;quot;alpha-* globin&amp;quot;) AND (taxonomy_id:8932)&amp;lt;/tt&amp;gt;, 2 hits.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=975</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=975"/>
		<updated>2026-09-02T13:41:25Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Simple text mining */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt homepage https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just containing the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&amp;lt;!-- This can be done very easily: To the left of the results list under &amp;lt;u&amp;gt;Search terms&amp;lt;/u&amp;gt; you find a list of links that allow you to restrict the search to specific fields (you may have to scroll down a bit). --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Model organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
&amp;lt;!-- *Under &amp;lt;u&amp;gt;Filter &amp;quot;human&amp;quot; as:&amp;lt;/u&amp;gt; click on: &amp;lt;u&amp;gt;organism&amp;lt;/u&amp;gt;. --&amp;gt;&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- *Under &amp;lt;u&amp;gt;Filter &amp;quot;insulin&amp;quot; as:&amp;lt;/u&amp;gt; click on: &amp;lt;u&amp;gt;protein name&amp;lt;/u&amp;gt;.   --&amp;gt;&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Note that all selections made with the mouse are shown in text format in the search field at the top of the page. It is possible to edit the search criteria manually in this field to make them broader or more narrow. &lt;br /&gt;
*Try for instance to exclude proteins that are not insulin, but only insulin-like. You do this by adding the following text in the search field: &#039;&#039;&#039;&amp;lt;tt&amp;gt;NOT name:insulin-like&amp;lt;/tt&amp;gt;&#039;&#039;&#039; and click on the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button. &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advaced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the actual cross-links to PDB. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. We will work with the American version of PDB later in the course. As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined several times. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Any experimental assertion&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of entries in Swiss-Prot may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;&amp;lt;u&amp;gt;or expand search to &amp;quot;1423&amp;quot; to include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=974</id>
		<title>Exercise: The protein database UniProt</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_The_protein_database_UniProt&amp;diff=974"/>
		<updated>2026-09-02T13:40:40Z</updated>

		<summary type="html">&lt;p&gt;Henni: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: Henrik Nielsen - updated by Morten Nielsen and Rasmus Wernersson&lt;br /&gt;
&lt;br /&gt;
[[File:Uniprotlogo.gif|right|frame|UniProt logo as of 2010 - source: http://www.uniprot.org]]&lt;br /&gt;
__TOC__&lt;br /&gt;
In this exercise, we shall extract information from the protein database, Uniprot. This database is administrated in collaboration between [https://www.sib.swiss/ Swiss Institute of Bioinformatics (SIB)], [http://www.ebi.ac.uk/ European Bioinformatics Institute (EBI)], England, and [http://www.georgetown.edu/ Georgetown University], Washington DC, USA.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
UniProt, http://www.uniprot.org/,  consists of three parts:&lt;br /&gt;
*&#039;&#039;&#039;UniProt Knowledge-base&#039;&#039;&#039; (UniProtKB) &lt;br /&gt;
**protein sequences with annotation and references&lt;br /&gt;
*&#039;&#039;&#039;UniProt Reference Clusters&#039;&#039;&#039; (UniRef) &lt;br /&gt;
**homology-reduced database, where similar sequences (having a certain percentage identity) are merged into clusters, each with a representative sequence &lt;br /&gt;
*&#039;&#039;&#039;UniProt Archive&#039;&#039;&#039; (UniParc) &lt;br /&gt;
**an archive containing all versions of Uniprot without annotations&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]]&lt;br /&gt;
Of these databases, &#039;&#039;&#039;Uniprot Knowledge-base is the most useful&#039;&#039;&#039;, and this is the database we shall be using today. Uniprot Knowledge-base consists of two parts:&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;Swiss-Prot&#039;&#039;&#039; &lt;br /&gt;
**a manually annotated (reviewed) protein-database.&lt;br /&gt;
*UniProtKB/&#039;&#039;&#039;TrEMBL&#039;&#039;&#039;&lt;br /&gt;
**a computer-annotated supplement to Swiss-Prot, that contains all translations of EMBL nucleotide sequences not yet included in Swiss-Prot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Simple text mining==&lt;br /&gt;
First, we will find some UniProt entries using simple text mining. You are supposed to find the entry for human insulin.&lt;br /&gt;
&lt;br /&gt;
*Open the UniProt home-page https://www.uniprot.org/&lt;br /&gt;
*Type &#039;&#039;&#039;&amp;lt;tt&amp;gt;human insulin&amp;lt;/tt&amp;gt;&#039;&#039;&#039; in the search field in the top of the page. Leave the search menu on &amp;quot;&amp;lt;u&amp;gt;UniProtKB&amp;lt;/u&amp;gt;&amp;quot;, which is default. Press Enter or click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
*If you are new to UniProt, you will be asked whether you want to view your results as &amp;quot;Cards&amp;quot; or &amp;quot;Table&amp;quot;. Choose &amp;quot;Table&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.1:&#039;&#039;&#039;&lt;br /&gt;
:#How many hits do you find? (tip: See the number above the results list)&lt;br /&gt;
:#How many of these hits are from Swiss-Prot? (tip: See under &amp;quot;&amp;lt;u&amp;gt;Reviewed&amp;lt;/u&amp;gt;&amp;quot; at the top left)&lt;br /&gt;
:#Can you identify the correct hit (&#039;&#039;i.e.&#039;&#039; see which one is actually human insulin and not something else)? If yes, write down is Accession code and Entry name (also called ID).&lt;br /&gt;
&lt;br /&gt;
In this case, it was relatively easy to spot the correct hit, but sometimes it is more difficult. If you do not identify the correct hit immediately, it will often help to narrow down the search, and that is exactly what we ask you to do in the next four questions. &lt;br /&gt;
&lt;br /&gt;
The first step is searching for proteins that actually come from the &#039;&#039;organism&#039;&#039; &amp;quot;human&amp;quot; and are &#039;&#039;named&#039;&#039; something containing the word &amp;quot;insulin&amp;quot;, as opposed to just containing the words &amp;quot;human&amp;quot; and &amp;quot;insulin&amp;quot; somewhere in the entry. &lt;br /&gt;
&amp;lt;!-- This can be done very easily: To the left of the results list under &amp;lt;u&amp;gt;Search terms&amp;lt;/u&amp;gt; you find a list of links that allow you to restrict the search to specific fields (you may have to scroll down a bit). --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
On the left, you can see a list of &amp;quot;Model organisms&amp;quot;. Try to click &amp;quot;Human&amp;quot;.&lt;br /&gt;
&amp;lt;!-- *Under &amp;lt;u&amp;gt;Filter &amp;quot;human&amp;quot; as:&amp;lt;/u&amp;gt; click on: &amp;lt;u&amp;gt;organism&amp;lt;/u&amp;gt;. --&amp;gt;&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.2:&#039;&#039;&#039; &lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- *Under &amp;lt;u&amp;gt;Filter &amp;quot;insulin&amp;quot; as:&amp;lt;/u&amp;gt; click on: &amp;lt;u&amp;gt;protein name&amp;lt;/u&amp;gt;.   --&amp;gt;&lt;br /&gt;
However, to really solve the problem, we have to enter Advanced mode. Click on &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; in the right part of the search field. Search for &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field, then click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; and search for &amp;lt;tt&amp;gt;insulin&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Protein Name [DE]&amp;lt;/u&amp;gt; field.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.3:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what has the search string in the text box at the top of the page now turned into? &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Note that all selections made with the mouse are shown in text format in the search field at the top of the page. It is possible to edit the search criteria manually in this field to make them broader or more narrow. &lt;br /&gt;
*Try for instance to exclude proteins that are not insulin, but only insulin-like. You do this by adding the following text in the search field: &#039;&#039;&#039;&amp;lt;tt&amp;gt;NOT name:insulin-like&amp;lt;/tt&amp;gt;&#039;&#039;&#039; and click on the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button. &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
Now, you should exclude proteins that are not insulin, but only insulin-like. Open the &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; menu again, add a field, make sure it is combined by &amp;lt;u&amp;gt;NOT&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, and remove hits that have &amp;lt;tt&amp;gt;insulin-like&amp;lt;/tt&amp;gt; in the protein name.&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.4:&#039;&#039;&#039;&lt;br /&gt;
:How many hits are now left? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
Note that you can also edit the search string directly, instead of going through the Advaced menu every time.&lt;br /&gt;
*Try now to exclude proteins that are insulin receptors (or substrates for insulin receptors). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 1.5:&#039;&#039;&#039; &lt;br /&gt;
:#How did you do this? &lt;br /&gt;
:#How many hits are now left? How many of these are from Swiss-Prot?&lt;br /&gt;
&lt;br /&gt;
==The contents of UniProt==&lt;br /&gt;
&lt;br /&gt;
We shall now see what information is contained in a UniProt entry, and what further information is available as links in each entry.&lt;br /&gt;
&lt;br /&gt;
Click on the accession-code or ID for insulin. This will take you to the insulin entry in the UniProtKB/Swiss-Prot database. Spend some time to get an overview of the page and the information it contains.&lt;br /&gt;
&lt;br /&gt;
*Note that you can click on the headings in the left side of the page to scroll to different sections of the page. Try it!&lt;br /&gt;
&lt;br /&gt;
*Note also that every time there is a small &amp;quot;&amp;lt;u&amp;gt;i&amp;lt;/u&amp;gt;&amp;quot; after a term on the page, you can click it to get information about the term. Try it!&lt;br /&gt;
&lt;br /&gt;
Now click on &amp;lt;u&amp;gt;Publications&amp;lt;/u&amp;gt; in the top part of the window. Click on &amp;lt;u&amp;gt;UniProtKB/Swiss-Prot&amp;lt;/u&amp;gt; under &amp;lt;u&amp;gt;Source&amp;lt;/u&amp;gt; to show only those references that are part of the entry and exclude those that are &amp;quot;computationally mapped&amp;quot;. Note that it is indicated what each reference has contributed (&amp;quot;&amp;lt;u&amp;gt;Cited for&amp;lt;/u&amp;gt;&amp;quot;). You can get to the PubMed literature database at NCBI by clicking at the link &amp;quot;&amp;lt;u&amp;gt;PubMed&amp;lt;/u&amp;gt;&amp;quot; for a reference — try this. The abstract of a publication can be read here (or directly in UniProt using the &amp;quot;&amp;lt;u&amp;gt;View abstract&amp;lt;/u&amp;gt;&amp;quot;-link), if the work is an actual published article and not a &amp;quot;direct submission&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.1:&#039;&#039;&#039; &lt;br /&gt;
:#How many references are there in the insulin entry? &lt;br /&gt;
:#Why do you think insulin is such a highly investigated protein? (Hint: see other sections of the entry, &#039;&#039;e.g.&#039;&#039; &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, especially the subsections &amp;lt;u&amp;gt;Involvement in disease&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Pharmaceutical&amp;lt;/u&amp;gt;)&lt;br /&gt;
&lt;br /&gt;
*Scroll back to &amp;lt;u&amp;gt;Function&amp;lt;/u&amp;gt; and read the free-text description at the top of the section. Also have a look at the controlled vocabulary annotations: &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt;. Note that both of these are split into two different aspects: &amp;lt;u&amp;gt;Molecular function&amp;lt;/u&amp;gt; and &amp;lt;u&amp;gt;Biological process&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Now scroll to &amp;lt;u&amp;gt;Subcellular Location&amp;lt;/u&amp;gt; and read what is written there. Note that you find another set of &amp;quot;Gene Ontology&amp;quot; (&amp;lt;u&amp;gt;GO&amp;lt;/u&amp;gt;) and &amp;lt;u&amp;gt;Keywords&amp;lt;/u&amp;gt; annotations here; this time labelled &amp;lt;u&amp;gt;Cellular component&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.2:&#039;&#039;&#039; &lt;br /&gt;
:#Where in the cell / outside the cell do you find insulin? &lt;br /&gt;
:#Why do you think is it found there? (Hint: consider the function)&lt;br /&gt;
&lt;br /&gt;
Just like in GenBank, a UniProt entry has a &#039;&#039;Feature Table&#039;&#039; containing annotations that are coupled to specific parts of the sequence. In the default view, the Feature Table is not so easy to spot, since it is split up under different sections corresponding to the biological significance of the various annotations. However, in the top part of the window you can click on &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt;, which shows the feature table information in a graphical form. Try it. Then click on &amp;lt;u&amp;gt;Molecule processing&amp;lt;/u&amp;gt; to show the signal peptide and the propeptide.&lt;br /&gt;
&lt;br /&gt;
Now switch back to the default (&amp;lt;u&amp;gt;Entry&amp;lt;/u&amp;gt;) view. In the following, you will see some examples of Feature Table annotations.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Disease &amp;amp; Drugs&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Variants&amp;lt;/u&amp;gt; lists the variants (mutations) of insulin that have been described in the literature. Under the heading &amp;lt;u&amp;gt;Change&amp;lt;/u&amp;gt;, it is indicated which amino acid is changed into which other amino acid. If the variant is known to be associated with a disease, this is indicated under the heading &amp;lt;u&amp;gt;Description&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
*Under &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows that insulin has both a signal peptide and a pro-peptide. These are both cleaved off before secretion. The mature insulin (the A and B chains) is hence much smaller than what was shown under &amp;lt;u&amp;gt;Sequences&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.3:&#039;&#039;&#039;&lt;br /&gt;
:How long is the signal peptide and the propeptide, respectively?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
*Under &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, the subsection &amp;lt;u&amp;gt;Features&amp;lt;/u&amp;gt; shows the secondary structure elements &amp;quot;Helix&amp;quot; (&amp;amp;alpha;-helix), &amp;quot;Beta strand&amp;quot; (part of a &amp;amp;beta;-pleated sheet) or &amp;quot;Turn&amp;quot;. The regions without specified secondary structure are often called &amp;quot;Loop&amp;quot; or &amp;quot;Coil&amp;quot;. &#039;&#039;&#039;CORRECTION 2024&#039;&#039;&#039;: With the latest update of the UniProt interface, you need to go to the top of the window and select &amp;lt;u&amp;gt;Feature viewer&amp;lt;/u&amp;gt; to see the secondary structure annotations! Click &amp;lt;u&amp;gt;Structural features&amp;lt;/u&amp;gt; on the left to see helices, strands, and turns. Click each coloured box to see positions.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2.4:&#039;&#039;&#039;&lt;br /&gt;
:Which positions are in &amp;amp;beta;-sheet conformation in insulin?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Other databases linked from UniProt===&lt;br /&gt;
&lt;br /&gt;
UniProt has many useful links to other databases. In the graphical view, the cross-references are spread among several different headings, just like the feature table is. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Sequence &amp;amp; Isoform&amp;lt;/u&amp;gt;, there is a sub-heading named &amp;lt;u&amp;gt;Sequence databases&amp;lt;/u&amp;gt;. Here, you can e,g, find links to nucleotide sequences in the databases EMBL / GenBank / DDBJ. Try clicking one of the GenBank links marked &amp;quot;Genomic DNA&amp;quot;; that should take you to a page that looks like something you have seen last week.&lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Structure&amp;lt;/u&amp;gt;, there is an interactive window showing a three-dimensional structure of insulin. Note that you can rotate the structure with your mouse. Actually, this structure is not part of UniProt itself, it is a cross-link to the protein structure database PDB. Below the interactive window, you can see the actual cross-links to PDB. Note that PDB is not one single database – just like it was the case for the nucleotide databases, there is a European version (PDBe), an American version (RCSB-PDB), and a Japanese version (PDBj), but luckily, they contain the same data. We will work with the American version of PDB later in the course. As you can see, there are many PDB structures of insulin; in other words, the 3D structure of insulin has been determined several times. &lt;br /&gt;
&lt;br /&gt;
Under the heading &amp;lt;u&amp;gt;Family &amp;amp; Domains&amp;lt;/u&amp;gt;, there is a subsection named &amp;lt;u&amp;gt;Family and domain databases&amp;lt;/u&amp;gt;. It has links to databases containing proteins that are similar (protein families). These have been collected using various techniques that you will hear about later in the course (multiple alignment). In some cases, the proteins are similar only in smaller parts (domains) but not in other parts, and in some cases the databases can tell which parts of the actual protein are known in other species. Some large proteins (not small ones like insulin) can contain several different parts (domains) each with their own evolutionary history. The most important of these databases is InterPro, because it collects the information from most of the other databases. Try to click on one of the InterPro links. This will take you to the Interpro page with lots of information about the protein family that insulin belongs to.&lt;br /&gt;
&lt;br /&gt;
===Text format===&lt;br /&gt;
&lt;br /&gt;
Until now, we have been working with the graphical user interface to UniProt. However, all the information is also available in plain text format, and that&#039;s what you will be working with if you are going to analyze larger amounts of UniProt data later in your studies. For now, let&#039;s just have a look at it. &lt;br /&gt;
&lt;br /&gt;
Scroll to the top of the Human Insulin page and find the menu labeled &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt;. It looks like this: [[File:Download.png]]. Click it, and then &#039;&#039;right-click&#039;&#039; the option &amp;lt;u&amp;gt;Text&amp;lt;/u&amp;gt; and open it in a new tab. What you see here basically contains all the information you have seen in the graphical interface.&lt;br /&gt;
&lt;br /&gt;
Scroll through the plain text file and see if you can find the same information that you just found in the graphical interface. Note that every line starts with a two-letter code specifying the type of the information in the line. Here are some examples:&lt;br /&gt;
* &#039;&#039;&#039;ID&#039;&#039;&#039;: Entry name (ID). There is only one ID.&lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039;: Accession code. There may be more than one.&lt;br /&gt;
* &#039;&#039;&#039;DE&#039;&#039;&#039;: Description (protein names).&lt;br /&gt;
* &#039;&#039;&#039;GN&#039;&#039;&#039;: Gene Name&lt;br /&gt;
* &#039;&#039;&#039;OS&#039;&#039;&#039;: Organism/Species.&lt;br /&gt;
* &#039;&#039;&#039;OC&#039;&#039;&#039;: Organism Classification.&lt;br /&gt;
* &#039;&#039;&#039;OX&#039;&#039;&#039;: TaxID (as defined in the NCBI Taxonomy database).&lt;br /&gt;
* &#039;&#039;&#039;RN&#039;&#039;&#039;, &#039;&#039;&#039;RP&#039;&#039;&#039;, &#039;&#039;&#039;RX&#039;&#039;&#039;, &#039;&#039;&#039;RA&#039;&#039;&#039;, &#039;&#039;&#039;RT&#039;&#039;&#039;, &#039;&#039;&#039;RL&#039;&#039;&#039;: References.&lt;br /&gt;
* &#039;&#039;&#039;CC&#039;&#039;&#039;: Comments (annotations pertaining to the whole protein).&lt;br /&gt;
* &#039;&#039;&#039;DR&#039;&#039;&#039;: Cross-references to other databases.&lt;br /&gt;
* &#039;&#039;&#039;KW&#039;&#039;&#039;: Keywords.&lt;br /&gt;
* &#039;&#039;&#039;FT&#039;&#039;&#039;: Feature Table (annotations pertaining to specified parts of the sequence).&lt;br /&gt;
* &#039;&#039;&#039;SQ&#039;&#039;&#039;: Sequence header line.&lt;br /&gt;
&lt;br /&gt;
==Advanced search==&lt;br /&gt;
&lt;br /&gt;
The UniProt interface allows you to use most of the fields in the database for searching, not only the fields like name and organism, as we did previously, but also the functional and structural annotations. We shall now try a few of these. &lt;br /&gt;
&lt;br /&gt;
* Go back to UniProt&#039;s main page, http://www.uniprot.org/. &amp;lt;!-- Go back to the main page of UniProt&#039;s beta website, https://beta.uniprot.org/ .--&amp;gt; &#039;&#039;&#039;Important:&#039;&#039;&#039; If the search string from the previous search is still shown in the search field, clear it. Then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; to the right of the search field. This brings up a box with a new interface.&lt;br /&gt;
&lt;br /&gt;
* Now we will find out how many proteins have signal peptides (just like insulin has). In the drop-down menu that appears in the box, select &amp;lt;u&amp;gt;PTM/Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Molecule Processing&amp;lt;/u&amp;gt;, then select &amp;lt;u&amp;gt;Signal peptide&amp;lt;/u&amp;gt;. In the empty field that now appears to the right of the word &amp;lt;u&amp;gt;Signal&amp;lt;/u&amp;gt;, type a &amp;lt;tt&amp;gt;*&amp;lt;/tt&amp;gt; (otherwise, it will not work). Click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.1:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins did you find, how many of them are from Swiss-Prot, and what was the search string (the text that appeared in the search field)?&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Evidence:&#039;&#039; The proteins we find in this way include proteins that are &#039;&#039;predicted&#039;&#039; to have signal peptides, without necessarily having any experimental evidence for the signal peptides. We will now limit the search to experimentally confirmed signal peptides. Click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again (without erasing your previous search) and change the &amp;lt;u&amp;gt;Evidence&amp;lt;/u&amp;gt; menu to &amp;lt;u&amp;gt;Any experimental assertion&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.2:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, how many of them are from Swiss-Prot, and what has the search string changed into? &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Combining fields:&#039;&#039; How many experimentally confirmed signal peptides are found in humans? Click on &amp;lt;u&amp;gt;Advanced Search&amp;lt;/u&amp;gt; again and click &amp;lt;u&amp;gt;Add field&amp;lt;/u&amp;gt; to get a second search line. Leave the menu to the left on &amp;lt;u&amp;gt;AND&amp;lt;/u&amp;gt;, select &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; in the drop-down menu, type &amp;lt;tt&amp;gt;human&amp;lt;/tt&amp;gt; in the field &amp;lt;u&amp;gt;Term&amp;lt;/u&amp;gt;, accept the suggestion &amp;quot;Homo sapiens (Human) [9606]&amp;quot; and click the &amp;lt;u&amp;gt;Search&amp;lt;/u&amp;gt; button.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.3:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins do you find now, and what is the search string? (Note that you can always perform the search by editing the text in the search field — however to do this you need to know the names for the fields).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;Important note&#039;&#039;&#039; about the organism field: when you type some letters, a drop-down list with suggestions will come up. Each has a number in brackets — this is the TaxID, which you can also find in [http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi the NCBI Taxonomy Browser]. If you search for &#039;&#039;e.g.&#039;&#039; Human proteins, it is a good idea to include the TaxID; if you omit it and just write &amp;quot;human&amp;quot;, you will also find proteins from organisms like Human immunodeficiency virus (try it!). &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== About strains and subspecies ===&lt;br /&gt;
Let us now try something different. If you search for proteins from a microbial species, you may run into trouble, because each subspecies or strain has its own TaxID, and you probably want all possible strains. Let&#039;s try an example (&#039;&#039;&#039;first, clear the previous search&#039;&#039;&#039;): Say you want all proteins from the bacterium &#039;&#039;Bacillus subtilis&#039;&#039; — a very important production organism in biotechnology. Try to type &amp;lt;tt&amp;gt;Bacillus subtilis&amp;lt;/tt&amp;gt; in the &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; field: you will see a suggestion named &amp;quot;Bacillus subtilis [1423]&amp;quot; – accept that.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.4:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; with the default TaxID [1423]? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Now note the line above the results that says &amp;quot;Expand search &amp;quot;Bacillus subtilis [1423]&amp;quot; to include lower taxonomic ranks&amp;quot; – click it. --&amp;gt;&lt;br /&gt;
The number of entries in Swiss-Prot may seem low for such a well-studied organism. In addition, you may note that there is a link next to the total number of results saying &amp;quot;&amp;lt;u&amp;gt;or expand search to &amp;quot;1423&amp;quot; to include lower taxonomic ranks&amp;lt;/u&amp;gt;&amp;quot;. Click it. &lt;br /&gt;
&amp;lt;!-- Now do the same thing, but in the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.5:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are there in UniProt from &#039;&#039;Bacillus subtilis&#039;&#039; in total (all strains and subspecies)? How many of these are from Swiss-Prot? And what is the search string?&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎|left]] In conclusion, use the field &amp;lt;u&amp;gt;Taxonomy [OC]&amp;lt;/u&amp;gt; instead of &amp;lt;u&amp;gt;Organism [OS]&amp;lt;/u&amp;gt; when working with microbial species where you want all strains.&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&lt;br /&gt;
&lt;br /&gt;
===Searching for short proteins===&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;Numerical field:&#039;&#039; Now we will try to answer a completely different question: Which extremely short proteins are present in UniProt? Clear the previous search. In the advanced drop-down menu, select &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt; and then &amp;lt;u&amp;gt;Sequence length&amp;lt;/u&amp;gt;. Now two new fields appear where you can define the lower and upper limits for the search. Type &amp;lt;tt&amp;gt;1&amp;lt;/tt&amp;gt; and &amp;lt;tt&amp;gt;10&amp;lt;/tt&amp;gt; and search. &#039;&#039;&#039;Note:&#039;&#039;&#039; in your answers to the questions below, include the search string just like you did in the questions above!&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.6:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins of maximum length 10 do you find?&lt;br /&gt;
&lt;br /&gt;
*Extremely short proteins are often mistakes translated directly from a nucleotide sequence with no evidence for the sequences being protein coding. Limit your search to proteins that actually have evidence for their existence at the protein level (add a field, and set the drop-down menu to &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; and select &amp;lt;u&amp;gt;Evidence at protein level&amp;lt;/u&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.7:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&amp;lt;blockquote style=&amp;quot;background-color: lightyellow; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;NOTE (February 2020):&#039;&#039;&#039; UniProt currently has a bug related to the &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; field. When you have made a search using &amp;lt;u&amp;gt;Protein existence&amp;lt;/u&amp;gt; and then click &amp;lt;u&amp;gt;Advanced&amp;lt;/u&amp;gt; again, it will convert the search to a search for &amp;quot;Evidence at protein level [1]&amp;quot; in &amp;lt;u&amp;gt;All&amp;lt;/u&amp;gt; fields. This will produce unexpected results. Therefore, you have to convert it &#039;&#039;manually&#039;&#039; back to a &amp;lt;u&amp;gt;Protein existence [PE]&amp;lt;/u&amp;gt; search before adding another criterion.&amp;lt;br&amp;gt; The bug has been reported to UniProt.&lt;br /&gt;
&amp;lt;/blockquote&amp;gt; &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
*A large fraction of the proteins identified in this way are fragments. Try to exclude fragments from the search. Add a field. In the drop-down menu, choose &amp;lt;u&amp;gt;Sequence&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;Fragment&amp;lt;/u&amp;gt;, then &amp;lt;u&amp;gt;No&amp;lt;/u&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.8:&#039;&#039;&#039; &lt;br /&gt;
:How many proteins are now left?&lt;br /&gt;
&lt;br /&gt;
*And how many of these proteins are found in humans?. Do as before...&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.9:&#039;&#039;&#039; &lt;br /&gt;
:How many human non-fragment proteins of maximum length 10 do you find in UniProt?&lt;br /&gt;
&lt;br /&gt;
*Finally you can save the results of your search. First, sort them by length by clicking on the column header. Then, click on &amp;lt;u&amp;gt;Download&amp;lt;/u&amp;gt; above the list of results. You can now save the search results in the format you prefer (try &amp;lt;u&amp;gt;FASTA (canonical)&amp;lt;/u&amp;gt; and click &amp;lt;u&amp;gt;Preview&amp;lt;/u&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3.10:&#039;&#039;&#039; &lt;br /&gt;
:Copy the FASTA sequences to your report.&lt;br /&gt;
&lt;br /&gt;
== On your own ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]] &#039;&#039;&#039;QUESTION 4&#039;&#039;&#039;: Now that you are proficient in UniProt searches, try the following:&lt;br /&gt;
&lt;br /&gt;
(As always, remember to write your search string in the answer).&lt;br /&gt;
# Find out how many proteins from &#039;&#039;Escherichia coli&#039;&#039; (all strains) there are in UniProt.&lt;br /&gt;
# How many of these are from the notorious pathogenic serotype [https://en.wikipedia.org/wiki/Escherichia_coli_O157:H7 O157:H7] (including its sub-strains)?&lt;br /&gt;
# Find insulin from as many organisms as possible, without including entries that are not insulin (&#039;&#039;&#039;Hint&#039;&#039;&#039;: If you attempt to do this with the Protein Name field only, it will require an unwieldy amount of kill-words. Therefore, take the gene name into account).&lt;br /&gt;
# Find alpha-globin (the alpha subunit of hemoglobin) from as many ruminants as possible (see the GenBank exercise).&lt;br /&gt;
# Find alpha-A globin and alpha-D globin from &#039;&#039;Columba livia&#039;&#039; (&#039;&#039;&#039;Hint&#039;&#039;&#039;: You can use a &amp;quot;*&amp;quot; to perform the search with one search string).&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_Translation_-_Virtual_Ribosome&amp;diff=973</id>
		<title>Exercise: Translation - Virtual Ribosome</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Exercise:_Translation_-_Virtual_Ribosome&amp;diff=973"/>
		<updated>2026-09-02T13:35:57Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Step 2: Genetic codes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In this exercise we will be using &#039;&#039;Virtual Ribosome&#039;&#039; - a software that provides a series of functions to &#039;&#039;&#039;translate DNA to protein sequences&#039;&#039;&#039;. Besides using the simple functions to translate DNA using a known reading frame, we shall work on computer-based analysis of possible reading frames, location of START and STOP codons etc.&lt;br /&gt;
&lt;br /&gt;
==Step 1: Basic translation==&lt;br /&gt;
* Open Virtual Ribosome (in a new window): https://services.healthtech.dtu.dk/services/VirtualRibosome-2.0/ Spend a few minutes to get familiar with the website - where do you upload the input data, and what types of options are available.&lt;br /&gt;
&lt;br /&gt;
* If you only have one sequence, this can be directly pasted into the input window. Alternatively, Virtual Ribosome can handle a series of different input formats that allow for multiple sequence inputs (i.e. FASTA).&lt;br /&gt;
&lt;br /&gt;
* Lets first do a simple example, and make a translation of a known gene Actin (from Yeast). Copy the sequence below into the sequence field and press &amp;quot;submit query&amp;quot;, using default settings.&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;Yeast_ACT1&lt;br /&gt;
 ATGGATTCTGAGGTTGCTGCTTTGGTTATTGATAACGGTTCTGGTATGTGTAAAGCCGGT&lt;br /&gt;
 TTTGCCGGTGACGACGCTCCTCGTGCTGTCTTCCCATCTATCGTCGGTAGACCAAGACAC&lt;br /&gt;
 CAAGGTATCATGGTCGGTATGGGTCAAAAAGACTCCTACGTTGGTGATGAAGCTCAATCC&lt;br /&gt;
 AAGAGAGGTATCTTGACTTTACGTTACCCAATTGAACACGGTATTGTCACCAACTGGGAC&lt;br /&gt;
 GATATGGAAAAGATCTGGCATCATACCTTCTACAACGAATTGAGAGTTGCCCCAGAAGAA&lt;br /&gt;
 CACCCTGTTCTTTTGACTGAAGCTCCAATGAACCCTAAATCAAACAGAGAAAAGATGACT&lt;br /&gt;
 CAAATTATGTTTGAAACTTTCAACGTTCCAGCCTTCTACGTTTCCATCCAAGCCGTTTTG&lt;br /&gt;
 TCCTTGTACTCTTCCGGTAGAACTACTGGTATTGTTTTGGATTCCGGTGATGGTGTTACT&lt;br /&gt;
 CACGTCGTTCCAATTTACGCTGGTTTCTCTCTACCTCACGCCATTTTGAGAATCGATTTG&lt;br /&gt;
 GCCGGTAGAGATTTGACTGACTACTTGATGAAGATCTTGAGTGAACGTGGTTACTCTTTC&lt;br /&gt;
 TCCACCACTGCTGAAAGAGAAATTGTCCGTGACATCAAGGAAAAACTATGTTACGTCGCC&lt;br /&gt;
 TTGGACTTCGAACAAGAAATGCAAACCGCTGCTCAATCTTCTTCAATTGAAAAATCCTAC&lt;br /&gt;
 GAACTTCCAGATGGTCAAGTCATCACTATTGGTAACGAAAGATTCAGAGCCCCAGAAGCT&lt;br /&gt;
 TTGTTCCATCCTTCTGTTTTGGGTTTGGAATCTGCCGGTATTGACCAAACTACTTACAAC&lt;br /&gt;
 TCCATCATGAAGTGTGATGTCGATGTCCGTAAGGAATTATACGGTAACATCGTTATGTCC&lt;br /&gt;
 GGTGGTACCACCATGTTCCCAGGTATTGCCGAAAGAATGCAAAAGGAAATCACCGCTTTG&lt;br /&gt;
 GCTCCATCTTCCATGAAGGTCAAGATCATTGCTCCTCCAGAAAGAAAGTACTCCGTCTGG&lt;br /&gt;
 ATTGGTGGTTCTATCTTGGCTTCTTTGACTACCTTCCAACAAATGTGGATCTCAAAACAA&lt;br /&gt;
 GAATACGACGAAAGTGGTCCATCTATCGTTCACCACAAGTGTTTCTAA&lt;br /&gt;
&lt;br /&gt;
* Look at the result. Note that the output shows both the DNA, and protein sequences as well as information on START and STOP codons. You can click on &amp;quot;instructions&amp;quot; on both the main page and the results page for details on what is displayed. Note also that the &amp;quot;raw&amp;quot; protein sequence can be downloaded in FASTA format.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:* &#039;&#039;&#039;QUESTION 1:&#039;&#039;&#039;&lt;br /&gt;
:# How is a STOP codon displayed?&lt;br /&gt;
:# How is a START codon displayed?&lt;br /&gt;
:# Does a start-codon always code for Methionine (M)?&lt;br /&gt;
:# What is the difference between the two types of start codons?&lt;br /&gt;
&lt;br /&gt;
==Step 2: Genetic codes==&lt;br /&gt;
* We are now going to work with yet another gene from yeast. This time it is COX1 that codes for Cytochrome C OXidase, subunit 1 (for more information click here: [https://www.yeastgenome.org/locus/S000007260 COX1 - Saccharomyces Genome Database]). Note that this is a &#039;&#039;&#039;mitochondrial gene&#039;&#039;&#039;. Translate this gene using default settings.&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;Yeast_COX1 &lt;br /&gt;
 ATGGTACAAAGATGATTATATTCAACAAATGCAAAAGATATTGCAGTATTATATTTTATG&lt;br /&gt;
 TTAGCTATTTTTAGTGGTATGGCAGGAACAGCAATGTCTTTAATCATTAGATTAGAATTA&lt;br /&gt;
 GCTGCACCTGGTTCACAATATTTACATGGTAATTCACAATTATTTAATGTTTTAGTAGTT&lt;br /&gt;
 GGTCATGCTGTATTAATGATTTTCTTCTTAGTAATGCCTGCTTTAATTGGAGGTTTTGGT&lt;br /&gt;
 AACTATTTATTACCATTAATAATTGGAGCTACAGATACAGCATTTCCAAGAATTAATAAC&lt;br /&gt;
 ATTGCTTTTTGAGTATTACCTATGGGGTTAGTATGTTTAGTTACATCAACTTTAGTAGAA&lt;br /&gt;
 TCAGGTGCTGGTACAGGGTGAACTGTCTATCCACCATTATCATCTATTCAGGCACATTCA&lt;br /&gt;
 GGACCTAGTGTAGATTTAGCAATTTTTGCATTACATTTAACATCAATTTCATCATTATTA&lt;br /&gt;
 GGTGCTATTAATTTCATTGTAACAACATTAAATATGAGAACAAATGGTATGACAATGCAT&lt;br /&gt;
 AAATTACCATTATTTGTATGATCAATTTTCATTACAGCGTTCTTATTATTATTATCATTA&lt;br /&gt;
 CCTGTATTATCTGCTGGTATTACAATGTTATTATTAGATAGAAACTTCAATACTTCATTC&lt;br /&gt;
 TTTGAAGTATCAGGAGGTGGTGACCCAATCTTATACGAGCATTTATTTTGATTCTTTGGT&lt;br /&gt;
 CACCCTGAAGTATATATTTTAATTATTCCTGGATTTGGTATTATTTCACATGTAGTATCA&lt;br /&gt;
 ACATATTCTAAAAAACCTGTATTTGGTGAAATTTCAATGGTATATGCTATGGCTTCAATT&lt;br /&gt;
 GGATTATTAGGATTCTTAGTATGATCACATCATATGTATATTGTAGGATTAGATGCAGAT&lt;br /&gt;
 CTTAGAGCATATTTCCTATCTGCACTAATGATTATTGCAATTCCAACAGGAATTAAAATT&lt;br /&gt;
 TTCTCATGATTAGCTCTAATCCATGGTGGTTCAATTAGATTAGCACTACCTATGTTATAT&lt;br /&gt;
 GCAATTGCATTCTTATTCTTATTCACAATGGGTGGTTTAACTGGTGTTGCCTTAGCTAAC&lt;br /&gt;
 GCCTCATTAGATGTAGCATTCCACGATACTTACTACGTGGTGGGACATTTTCACTATGTA&lt;br /&gt;
 TTATCAATGGGTGCTATTTTCTCTTTATTTGCAGGATACTATTATTGAAGTCCTCAAATT&lt;br /&gt;
 TTAGGTTTAAACTATAATGAAAAATTAGCTCAAATTCAATTCTGATTAATTTTCATTGGG&lt;br /&gt;
 GCTAATGTTATTTTCTTCCCAATGCATTTTTTAGGTATTAATGGTATGCCTAGAAGAATT&lt;br /&gt;
 CCTGATTATCCTGATGCTTTCGCAGGATGAAATTATGTCGCTTCTATTGGTTCATTCATT&lt;br /&gt;
 GCACTATTATCATTATTCTTATTTATCTATATTTTATATGATCAATTAGTTAATGGATTA&lt;br /&gt;
 AACAATAAAGTTAATAATAAATCAGTTATTTATAATAAAGCACCTGATTTTGTAGAATCT&lt;br /&gt;
 AATCTTATCTTTAATTTAAATACAGTTAAATCTTCATCTATCGAATTCTTATTAACTTCT&lt;br /&gt;
 CCACCAGCTGTACACTCATTTAATACACCAGCTGTACAATCTTAA&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 2&#039;&#039;&#039;:&lt;br /&gt;
:# Did the translation succeed (&#039;&#039;i.e.&#039;&#039; did it yield a long amino acid sequence unbroken by stop codons)? &lt;br /&gt;
:# Nothing is wrong with the DNA sequence. Can you come up with some good reasons for the result?&lt;br /&gt;
&lt;br /&gt;
* Keep the result of the translation in a window (we need it again in a while), and open a new window with Virtual Ribosome. Translate the DNA sequence once more using a different translation table (see options). Guess yourself which table to select. &lt;br /&gt;
&lt;br /&gt;
* If you have chosen the right translation table, the DNA sequence can be translated without any problems. Compare the two results and answer the following questions:&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 3&#039;&#039;&#039;&lt;br /&gt;
:# What is the difference in the use of STOP codons?&lt;br /&gt;
:# What is the difference in the use of START codons?&lt;br /&gt;
:# Are codons coding for completely different amino acids?&lt;br /&gt;
&lt;br /&gt;
More information on the definition of the different translation tables is found here: [http://www.ncbi.nlm.nih.gov/Taxonomy/Utils/wprintgc.cgi?mode=c The Genetic Codes - NCBI]. The tables are shown in a &amp;quot;compressed&amp;quot; format, but can be shown in a more comprehensible format by using the &amp;quot;&amp;lt;u&amp;gt;Click here to change format&amp;lt;/u&amp;gt;&amp;quot; option. Note:&lt;br /&gt;
The use of START codons is described in details for all genetic codes.&lt;br /&gt;
&#039;&#039;The difference between the standard-code and other codes is summarized in each section&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
==Step 3: Reading frames==&lt;br /&gt;
&#039;&#039;&#039;Remember to reset all options (in particular make sure that you now use the standard genetic code) before continuing the exercise.&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
We have up to now assumed that the reading frame for the DNA-sequence was known and that it always started at the first nucleotide. In the following, we shall examine how it is often possible to identify the most likely reading frame using computational translation tools. We shall use the the sequence below which is &#039;&#039;&#039;the complete mRNA sequence&#039;&#039;&#039; for a yeast gene (profilin). Use your biological knowledge to answer the following questions:&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 4:&#039;&#039;&#039;&lt;br /&gt;
:# Yeast has introns in some genes, could this be a major problem in this case?&lt;br /&gt;
:# Can an mRNA molecule contain more sequence than the gene in question? (Can it be longer than the CDS coding for the protein).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 &amp;gt;gi|4226|emb|Y00469.1| Yeast mRNA for profilin&lt;br /&gt;
 GGCAAATTATGTCTTGGCAAGCATACACTGATAACTTAATAGGAACCGGTAAAGTCGACAAAGCTGTCAT&lt;br /&gt;
 CTACTCGAGAGCAGGTGACGCTGTTTGGGCTACTTCTGGTGGCCTATCTTTGCAACCAAACGAAATTGGT&lt;br /&gt;
 GAAATTGTTCAAGGCTTCGACAATCCAGCTGGTTTGCAAAGCAATGGTTTGCATATTCAAGGCCAAAAGT&lt;br /&gt;
 TCATGTTGTTGAGAGCTGACGATAGAAGTATCTACGGTAGACATGATGCTGAGGGTGTTGTTTGTGTAAG&lt;br /&gt;
 AACTAAGCAAACCGTTATTATTGCTCATTATCCACCAACCGTACAAGCCGGTGAGGCCACCAAGATTGTC&lt;br /&gt;
 GAGCAATTGGCTGACTACTTGATTGGTGTTCAATACTAATTTATGCAGGTAAAGTTTTCTTGCCTTATAC&lt;br /&gt;
 ACCACCTATTCTGGCATCTGCGGGATTTCGCTTCCTATTTTACAAATATTTTATTGATTGACGCTAATTA&lt;br /&gt;
 TCACTGTAAAAGGCGCACTTTTTATATGTAGTCACATCCGGTATTTAACATATTTACGAAACAGTCTTAA&lt;br /&gt;
 GAATATCGACATTTGATATACTTATGTTTAATTTATCTACATATTACAATCA&lt;br /&gt;
&lt;br /&gt;
Six reading frames exist: 1, 2, 3 (on the positive stand, i.e. the sequence as you read it), and -1, -2, -3 (on the negative strand, i.e the complementary DNA string). Since we are working with a mRNA sequence, we do not need to consider the reading frames on the complementary string.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 5&#039;&#039;&#039;: &lt;br /&gt;
:Why is this?&lt;br /&gt;
&lt;br /&gt;
Translate the mRNA sequence in the three positive reading frames (1, 2, 3). The easiest way to do this, is to use a window/tab for each translation to be able to compare the different results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 6&#039;&#039;&#039;:&lt;br /&gt;
:What reading frame is most likely the right one?&lt;br /&gt;
&lt;br /&gt;
NB: remember that START and STOP codons are only shown for the selected reading frame.&lt;br /&gt;
Note also that the DNA-sequence is shown unmodified in all three reading frames whereas the protein sequence is shifted. &lt;br /&gt;
&lt;br /&gt;
It is possible to show multiple reading frames simultaneously. Use the Plus (1,2,3) as reading frame, and translate the sequence again.&lt;br /&gt;
&lt;br /&gt;
Note that the amino acid letter is centered above each codon (&#039;&#039;i.e.&#039;&#039; &amp;quot;M&amp;quot; is placed over the &amp;quot;T&amp;quot; in &amp;quot;ATG&amp;quot;).&lt;br /&gt;
The translation from reading frame 1 is shown just above the DNA sequence, followed by reading frame 2, and 3.&lt;br /&gt;
START and STOP codons for all three reading frames are shown at once&lt;br /&gt;
&lt;br /&gt;
For the sake of illustration, we shall try to translate the sequence on the negative strand. Select reading frame -1, and redo the translation.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
:# How does the DNA sequence in the output look (is it identical to the one you input)? &lt;br /&gt;
:# In what direction shall it be read (left-to-right or right-to-left)?&lt;br /&gt;
:# In what direction shall the protein-sequence be read (left-to-right or right-to-left)? &lt;br /&gt;
:(Try to compare to the protein sequence in FASTA format).&lt;br /&gt;
&lt;br /&gt;
Now, lets try to do it all in one go. Select All (6 reading frames) and translate the sequence again.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
:# How many DNA strings are displayed? &lt;br /&gt;
:# Why is this?&lt;br /&gt;
&lt;br /&gt;
Note the large number of possibilities a single DNA sequence contains with respect to translation to protein sequence.&lt;br /&gt;
&lt;br /&gt;
==Step 4: ORF finder==&lt;br /&gt;
&lt;br /&gt;
We have now made a manual screening for possible reading frames. Such a procedure might work fine if you have only one DNA sequence, but this is in general not the case, and often you need to use computer-based &#039;&#039;ORF finders&#039;&#039;. An ORF (Open Reading Frame) is a DNA sequence that is &#039;&#039;not interrupted by a STOP codon&#039;&#039;. Often one will be looking for the longest ORF starting with a START codon and ending at a STOP codon.&lt;br /&gt;
&lt;br /&gt;
The longest ORF is found by translating the sequence in all six reading frames, and then selecting the longest protein sequence.&lt;br /&gt;
&lt;br /&gt;
We shall now use a build-in ORF finder with the most stringent criteria. &lt;br /&gt;
Under in the ORF finder section use the following settings:&lt;br /&gt;
*Start codon: strict (this forces the ORF to start at ATG)&lt;br /&gt;
* Select &amp;quot;All (6 reading frames)&amp;quot; &lt;br /&gt;
&lt;br /&gt;
Finally translate the sequence using these settings.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
:&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
:# Does the result fit to what you found earlier?&lt;br /&gt;
:# Would it make any difference to the result if we had only a partial sequence where the last part of the sequence with the STOP codon is missing?&lt;br /&gt;
:# What would happen if the first 50 nucleotides (with the START codon) were missing?&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=971</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=971"/>
		<updated>2026-08-31T15:04:25Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Sep 1 — Introduction, taxonomy, and GenBank */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;Aud. 53, building 208&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/johanne-badsberg-overgaard?id=147205&amp;amp;entity=profile Johanne Badsberg Overgaard] &amp;amp;mdash; PhD student, stand-in for Melanie in week 1&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in week 2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible when you have uploaded your hand-in. You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course and bioinformatics&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Reference databases: Taxonomy and DNA sequences&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=970</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=970"/>
		<updated>2026-08-31T13:34:23Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Sep 1 — Introduction, taxonomy, and GenBank */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;Aud. 53, building 208&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/johanne-badsberg-overgaard?id=147205&amp;amp;entity=profile Johanne Badsberg Overgaard] &amp;amp;mdash; PhD student, stand-in for Melanie in week 1&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in week 2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible when you have uploaded your hand-in. You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=922</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=922"/>
		<updated>2026-08-28T12:59:19Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Question 3: */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) — this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just means a group) — sits between kingdom (Metazoa) and phylum (&#039;&#039;Chordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly)).&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px|left]]&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
We can ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=921</id>
		<title>Using the Taxonomy database</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=921"/>
		<updated>2026-08-28T12:39:13Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Example: Homo sapiens */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: &#039;&#039;&#039;Rasmus Wernersson&#039;&#039;&#039; and &#039;&#039;&#039;Henrik Nielsen&#039;&#039;&#039;.&lt;br /&gt;
 &lt;br /&gt;
==Background==&lt;br /&gt;
When comparing DNA and protein sequences from different species it is important to keep in mind that all living organisms at some point in time has shared a common ancestor. Some organisms are closely related and have recently derived from a common ancestor (e.g. Human and Chimpanzee, which diverged 5-10 million years ago) and some are more distantly related (e.g. Human and mouse, diverged 100-150 million years ago).&lt;br /&gt;
&lt;br /&gt;
The more closely related two organisms are, the more similar their sequences will be (say, when comparing the Alpha Globin gene from each of the organisms), and the more likely it will be that similar looking genes from each organism still have the same function (MUCH more about this when we get to pairwise alignment and BLAST searches).&lt;br /&gt;
&lt;br /&gt;
[[file:ApesTax_600.png‎|center|frame|Apes taxonomy (Detailed taxonomy of the Great Apes: Human, Chimp, Gorilla, Orangutan - from Wikipedia)]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Phylogeny vs. taxonomy===&lt;br /&gt;
As we discussed in the lecture all life is organised in a hierarchical taxonomical system, which approximates the &amp;quot;true&amp;quot; underlying phylogeny to a large degree. It&#039;s therefore often important to know where a specific organism is placed in the taxonomical system - this type of information will also always be included along with DNA/Protein sequences from the big databases such as GenBank and UniProt.&lt;br /&gt;
&lt;br /&gt;
Today we will explore various ways to look up and compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
===A word about Wikipedia===&lt;br /&gt;
The free online encyclopedia [http://en.wikipedia.org/wiki/Main_Page Wikipedia] (and other similar resources) is a GREAT way to start out when you need to look up information about a new topic - in this case taxonomy. Almost all species entries in Wikipedia has a &amp;quot;Scientific classification&amp;quot; box which includes taxonomical information (for example see the entry on [http://en.wikipedia.org/wiki/Orangutan Orangutan] or [http://en.wikipedia.org/wiki/Fucus_vesiculosus Fucus vesiculosus] (Bladder wrack / Blæretang)).&lt;br /&gt;
&lt;br /&gt;
HOWEVER: Keep in mind that Wikipedia is NOT a reliable source of information, even if most entries are of a very good quality. The facts in the Wikipedia entries have not been verified by taxonomy experts and can potentially be wrong (everybody can go in and edit the text). We need to look up the taxonomy in an official database (in this case we&#039;ll be using NCBI Taxonomy) before you can state it as a fact.&lt;br /&gt;
&lt;br /&gt;
You CANNOT quote Wikipedia as the only source of your information - you&#039;ll need to find the original primary source of the information or look it up in an official database.&lt;br /&gt;
&lt;br /&gt;
===A word about AI===&lt;br /&gt;
As with Wikipedia using AI to generate a taxonomical analysis (e.g. comparing how a bunch of species are related) can be a great way to get an overview, and a (typically) well written explanation. You will need to ask the AI to include references &#039;&#039;&#039;to actual scientific sources&#039;&#039;&#039; to document the validity of the results, and &#039;&#039;&#039;you will be responsible&#039;&#039;&#039; for double checking the AI output and making sure not only the data is correct, but also that the conclusions are sound.&lt;br /&gt;
&lt;br /&gt;
===The need for a Ground Truth===&lt;br /&gt;
Luckily, modern taxonomy is very well established and there an internationally recognized organization that makes revisions in a highly regulated manner. This &#039;&#039;&#039;taxonomy standard&#039;&#039;&#039; is captured in the &#039;&#039;&#039;NCBI Taxonomy&#039;&#039;&#039; database we&#039;ll be working with in the sections below. For all molecular data NCBI Taxonomy is THE standard reference.&lt;br /&gt;
&lt;br /&gt;
==The NCBI Taxonomy Database==&lt;br /&gt;
[[File:NcbiTax2026.jpg|800px|center|border]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Main link:&#039;&#039;&#039; https://www.ncbi.nlm.nih.gov/datasets/taxonomy/&lt;br /&gt;
&lt;br /&gt;
(&#039;&#039;2026 note: NCBI is in process of migrating to this new site - if you find the old Ncbi Tax homepage via Google, make sure to press the link they provide to jump to the new site&#039;&#039;).&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
As mentioned above NCBI Taxonomy will serve as our &#039;&#039;&#039;Ground Truth&#039;&#039;&#039; for everything related to taxonomy. The NCBI Tax database provides the numerical enumeration of species (and other taxonomical levels) that is used and referenced in most Sequence databases, such as GenBank (DNA) and UniProt (Protein). For example human (&#039;&#039;Homo sapiens&#039;&#039;) has the ID &amp;quot;&#039;&#039;&#039;9606&#039;&#039;&#039;&amp;quot; and Yeast (&#039;&#039;Saccharomyces cerevisiae&#039;&#039;) as the ID &amp;quot;&#039;&#039;&#039;4932&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
NCBI Tax is perhaps not a database you would browse for fun (depending on your level of geekiness). It&#039;s good for looking up definitions, and for comparing the taxonomical position of multiple organisms (since the information is so densely presented).&lt;br /&gt;
&lt;br /&gt;
===Example: Homo sapiens===&lt;br /&gt;
# Open the NCBI Taxonomy webpage in a new browser window/tab (see link above)&lt;br /&gt;
# Search for &amp;quot;&#039;&#039;Homo sapiens&#039;&#039;&amp;quot;.&lt;br /&gt;
## Dont Panic: An enormous amount of information is shown - for example about genome sequences. In this case we only need to look at the information presented in the &amp;quot;&#039;&#039;&#039;Taxonomy&#039;&#039;&#039;&amp;quot; tab.&lt;br /&gt;
## Notice the Taxonomy ID - &#039;&#039;&#039;9606&#039;&#039;&#039; as mentioned above.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Lineage&#039;&#039;&#039;&amp;quot; (box on the right hand side): Here a condensed overview of the human lineage is shown by default. Notice the taxonomical ranks we talked about in the lecture (&amp;quot;Phylum&amp;quot;, &amp;quot;Class&amp;quot; etc). You can navigate to the definition of these groups by clicking on them.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Full lineage&#039;&#039;&#039;&amp;quot;: Click this to view a FULL list of all the groups leading &amp;quot;down&amp;quot; to human. Notice, that you can &amp;quot;mouse over&amp;quot; the groups to see a pop-up with the taxonomical rank. NOTICE: The rank &amp;quot;CLADE&amp;quot; is used whenever a group does not have a common English name (&amp;quot;Clade&amp;quot; is the generic name for a uniquely defined taxonomical group).&lt;br /&gt;
&lt;br /&gt;
Play around with the Homo sapiens page for a bit to familiarize yourself with the interface, and answer the following questions along the way:&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1a&#039;&#039;&#039;: What is the TaxID of &amp;quot;&#039;&#039;Metazoa&#039;&#039;&amp;quot;? &lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1b&#039;&#039;&#039;: What is the family that contains humans?&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1c&#039;&#039;&#039;: Are humans vertebrates? (Latin: Vertebrata)?&lt;br /&gt;
&lt;br /&gt;
===Comparing taxonomy using NCBI Tax===&lt;br /&gt;
[[Image:Fruitfly.jpg|thumb|400px|right|Fruit fly (&#039;&#039;Drosophila melanogaster&#039;&#039;) - source: [http://en.wikipedia.org/wiki/Drosophila_melanogaster Wikipedia] ]]Besides being useful for being the official database behind the TaxID&#039;s used in GenBank (and other databases), NCBI Tax actually makes it easy to compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
Let&#039;s take the situation where you have read an interesting paper comparing a DNA sequence between the following three organisms: &#039;&#039;Homo sapiens&#039;&#039; (Human), &#039;&#039;Mus musculus&#039;&#039; (Mouse), and &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly), but you have no idea about the relationship between the three organisms. &lt;br /&gt;
&lt;br /&gt;
We can look this up in NCBI Tax:&lt;br /&gt;
&lt;br /&gt;
# Open two browser windows/tabs and search for &#039;&#039;&#039;Homo sapiens&#039;&#039;&#039; and &#039;&#039;&#039;Mus musculus&#039;&#039;&#039;.&lt;br /&gt;
# By comparing the &amp;quot;lineage&amp;quot; text it will be easy to find out at which taxonomical level human and mouse differ.&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2A&#039;&#039;&#039;: Look at the (simple/abbreviated) lineage information and find lowest ranking common group for human and mouse - what is the name and what is the rank?&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2B&#039;&#039;&#039;: Look at the &amp;quot;&#039;&#039;&#039;full lineage&#039;&#039;&#039; and find the lowest ranking group human and mouse have in common (it&#039;s OK if the rank is &amp;quot;clade&amp;quot;). What is the name and TaxId of the group?&lt;br /&gt;
&lt;br /&gt;
Now repeat the analysis by comparing Human and Fruit Fly (&#039;&#039;Drosophila melanogaster&#039;&#039;)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3A&#039;&#039;&#039;: What is the lowest rank group shared in the (simplified/abbreviated) lineage? (TaxID, Name)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3B&#039;&#039;&#039;: What is the lowest rank group shared in the full lineage? (TaxID, Name)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Fishing in NCBI Tax using the Common Tree function===&lt;br /&gt;
[[Image:Zebrafisch.jpg|thumb|400px|right|Zebrafish (&#039;&#039;Danio rerio&#039;&#039;) - source: [https://en.wikipedia.org/wiki/Zebrafish Wikipedia] ]]                               &lt;br /&gt;
In this last part of the exercise, we will investigate relationships between different species of fish. We have compiled this list of various fish:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;overflow:auto;&amp;quot;&amp;gt;&lt;br /&gt;
 Latin name             Common name                     TaxID&lt;br /&gt;
 Danio rerio            Zebrafish                        7955 &lt;br /&gt;
 Gadus morhua           Atlantic cod                     8049 &lt;br /&gt;
 Mustelus griseus       Spotless smooth-hound (shark)   89020  &lt;br /&gt;
 Petromyzon marinus     Sea lamprey                      7757 &lt;br /&gt;
 Latimeria chalumnae    Coelacanth (famous &amp;quot;Blue fish&amp;quot;)  7897 &lt;br /&gt;
 Lepidosiren paradoxa   South American lungfish          7883&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
Comparing many species will be tedious by just pointing and clicking, so here we&#039;ll utilize the NCBI &#039;&#039;&#039;Common Tree&#039;&#039;&#039; tool. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Now, go to [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/ the front page of NCBI Taxonomy] and click the &#039;&#039;&#039;Common Tree&#039;&#039;&#039; link a bit down the page (under &amp;quot;Taxonomy resources&amp;quot;). Here, you can add species to a tree one by one by entering either the Latin name or the TaxID in the &amp;quot;&#039;&#039;&#039;search and select&#039;&#039;&#039;&amp;quot; field. (&#039;&#039;It will try to autocomplete the name you&#039;re entering, so pay attention to what is selected when you press enter&#039;&#039;). You can also add a whole list at once if you have a text file containing &#039;&#039;either&#039;&#039; TaxIDs &#039;&#039;or&#039;&#039; Latin names, one per line (doable with a bit of editing of the list above).&lt;br /&gt;
&lt;br /&gt;
[[Image:NcbiCommonTree2026.jpg |400px|frame|center|&#039;&#039;&#039;Direct link:&#039;&#039;&#039; [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/common-tree/ Common Tree tool (2026)] ]]    &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; The 2026 version of the tool is quite good at automatically detecting the shared groups to display, but you can ask the tool to display a full lineage, if needed: You can test this by clicking &amp;quot;&#039;&#039;&#039;View in Taxonomy Browser&#039;&#039;&#039;&amp;quot; after the Common Tree has been populated (but &#039;&#039;&#039;return to the standard view&#039;&#039;&#039; for answering the questions below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4a:&#039;&#039;&#039; In this selection of species, what is the sister group (nearest neighbour) to the Zebrafish? What is the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
Now try to add yourself (i.e. Human) to the tree, using either the Latin name or the TaxID. Any surprises?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4b:&#039;&#039;&#039; What is now the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4c:&#039;&#039;&#039; Which of the following is most closely related to the &amp;quot;Blue fish&amp;quot;: the cod, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4d:&#039;&#039;&#039; Which of the following is most closely related to the cod: the shark, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4e:&#039;&#039;&#039; Which of the following is most closely related to the shark: the lamprey, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4f:&#039;&#039;&#039; Does the category &amp;quot;fish&amp;quot; make any scientific sense?&lt;br /&gt;
&lt;br /&gt;
Leave the browser window with the Common tree open for the next question.&lt;br /&gt;
&lt;br /&gt;
===Comparing trees===&lt;br /&gt;
A bioinformatician has compared the sequences of a gene from the seven species we used in the previous question, and arrived at the following tree:&lt;br /&gt;
&lt;br /&gt;
[[Image:FishNCBI-edited-phy.png]]&lt;br /&gt;
&lt;br /&gt;
You will later learn how to make trees like these in the [[Exercise: Phylogeny|Phylogenetic trees exercise]]. For now, you only need to know that such a tree is not necessarily 100% correct, since it is based on a limited amount of data. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5a:&#039;&#039;&#039; Are there any differences in the branching pattern between the gene tree and the Common tree from the previous question?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5b:&#039;&#039;&#039; Can the gene tree be made to comply with the Common tree by swapping two species? If so, which two?&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
===AI investigation===&lt;br /&gt;
As the final step, use your favorite AI (ChatGPT etc) to try to get as &#039;&#039;&#039;detailed and comprehensive&#039;&#039;&#039; an answer to QUESTION 4f as possible: &#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We do realize that the AI field is rapidly evolving, and answers will vary. We will recommend you to play around with the prompt to make the system generate an answer with the best possible scientific rigor. You are welcome to either use the list of species (+human) as we did above, or ask the AI to find relevant data itself.&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;QUESTION 6:&#039;&#039;&#039; Provide your final prompt + answer from the AI (+ info about which version you used). Reflect upon the correctness of the answer (how much do you trust it, &#039;&#039;anything fishy?&#039;&#039;)&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=853</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=853"/>
		<updated>2026-08-27T08:41:30Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Hand-ins */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;Aud. 53, building 208&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/johanne-badsberg-overgaard?id=147205&amp;amp;entity=profile Johanne Badsberg Overgaard] &amp;amp;mdash; PhD student, stand-in for Melanie in week 1&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in week 2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible when you have uploaded your hand-in. You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=852</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=852"/>
		<updated>2026-08-27T07:27:39Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Teaching assistants */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;Aud. 53, building 208&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/johanne-badsberg-overgaard?id=147205&amp;amp;entity=profile Johanne Badsberg Overgaard] &amp;amp;mdash; PhD student, stand-in for Melanie in week 1&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in week 2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=851</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=851"/>
		<updated>2026-08-20T09:44:54Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Where and when */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;Aud. 53, building 208&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=850</id>
		<title>Checklist for computers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=850"/>
		<updated>2026-08-20T09:37:09Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Software */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;At the exam in Introduction to Bioinformatics, you are going to use the same resources (web-servers and programs) that you used for the exercises — if your computer has worked fine for all the exercises, it will also work fine for the exam.&lt;br /&gt;
&lt;br /&gt;
== Hardware ==&lt;br /&gt;
; A laptop.&lt;br /&gt;
: It makes no difference whether you use Windows, Mac, or Linux, as long as you have the listed software. However, an iPad, an Android tablet, or a Chromebook will NOT be enough for the exam.&lt;br /&gt;
; A mouse.&lt;br /&gt;
: A mouse is important in order to use PyMOL (see below) optimally. The mouse should have two buttons plus a scroll-wheel in the middle.&lt;br /&gt;
&lt;br /&gt;
== Internet connection ==&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;NOTE 2021:&#039;&#039;&#039; Due to the corona situation, your exam will be from home. Whether you use wireless or cabled connection is not important, but &#039;&#039;it is your own responsibility that your hardware (computer, power supply, internet connection) is working during the exam&#039;&#039;. There will not be extra time granted on the basis of individual connection problems.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Your computer must be able to connect to DTU wi-fi. &amp;lt;!-- &#039;&#039;&#039;Use the standard DTU net which will be open during your exam; NOT the special exam net.&#039;&#039;&#039; --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Software ==&lt;br /&gt;
; A modern Internet Browser ([http://www.mozilla.com/ FireFox], Edge, Safari (Mac only), [http://www.google.com/chrome Google Chrome], [http://www.opera.com/ Opera]).&lt;br /&gt;
: NB: You &#039;&#039;must&#039;&#039; have FireFox, Chrome, or Opera on your machine, so you are able to switch browser in case Edge (Windows) or Safari (Mac) has trouble with a certain website.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Java&lt;br /&gt;
: Download and install from https://www.oracle.com/java/technologies/downloads/ (free) or https://adoptium.net/ (free), in case it is not already on your system. &lt;br /&gt;
: Test that it works:&lt;br /&gt;
:* jEdit, Jalview, and FigTree (see below) must be able to run.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
; Geany (or another GOOD plain text editor)&lt;br /&gt;
: Download and install from http://geany.org/, see the [[Plain text files and Geany]] exercise.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
: If you have unsolvable problems with jEdit, a good alternative is [ Geany].&lt;br /&gt;
: You MUST be able to handle text files with Unix, Mac, or Windows line endings — therefore, Notepad under Windows is NOT good enough.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; PyMol&lt;br /&gt;
: Download and install from https://pymol.org/, see the [[Media:PyMOL_tutorial.pdf|PyMol tutorial]] and the exercises in [[Protein Structure and Visualization|PDB &amp;amp; PyMol]], and [[Exercise:Malaria Vaccine|Malaria vaccine]]. &amp;lt;!-- , and [[ExPSIBLAST|PSI-BLAST]]. --&amp;gt;&lt;br /&gt;
: A license file is found on Learn → Content → Week 06. &lt;br /&gt;
&lt;br /&gt;
; Seaview&lt;br /&gt;
: Download and install from http://doua.prabi.fr/software/seaview, see [[Exercise: Multiple Alignments (Seaview version)]].&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Jalview&lt;br /&gt;
: Download and install from https://www.jalview.org/, see the exercise in [[Exercise: Multiple Alignments (English version)|Multiple Alignments]]. &lt;br /&gt;
; FigTree&lt;br /&gt;
: Download and install from https://github.com/rambaut/figtree/releases, see [[Exercise: Phylogeny]].&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; Text processing software&lt;br /&gt;
: for writing your answers. You can e.g. use:&lt;br /&gt;
:* Microsoft Word &lt;br /&gt;
:* [http://www.openoffice.org/ OpenOffice] / [http://www.neooffice.org/ NeoOffice] / [http://www.libreoffice.org/ LibreOffice] &lt;br /&gt;
:* Pages (for Mac).&lt;br /&gt;
:* [https://docs.google.com/ Google Docs] &lt;br /&gt;
&lt;br /&gt;
; Tool for making PDF files&lt;br /&gt;
: included in Windows 10/11 and Mac. &lt;br /&gt;
&lt;br /&gt;
; Tool for taking screenshots&lt;br /&gt;
: included in Microsoft Word, Windows 10/11, and Mac. &lt;br /&gt;
&amp;lt;!-- : For Windows users we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots.--&amp;gt;&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=849</id>
		<title>Checklist for computers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=849"/>
		<updated>2026-08-20T09:31:55Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Software */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;At the exam in Introduction to Bioinformatics, you are going to use the same resources (web-servers and programs) that you used for the exercises — if your computer has worked fine for all the exercises, it will also work fine for the exam.&lt;br /&gt;
&lt;br /&gt;
== Hardware ==&lt;br /&gt;
; A laptop.&lt;br /&gt;
: It makes no difference whether you use Windows, Mac, or Linux, as long as you have the listed software. However, an iPad, an Android tablet, or a Chromebook will NOT be enough for the exam.&lt;br /&gt;
; A mouse.&lt;br /&gt;
: A mouse is important in order to use PyMOL (see below) optimally. The mouse should have two buttons plus a scroll-wheel in the middle.&lt;br /&gt;
&lt;br /&gt;
== Internet connection ==&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;NOTE 2021:&#039;&#039;&#039; Due to the corona situation, your exam will be from home. Whether you use wireless or cabled connection is not important, but &#039;&#039;it is your own responsibility that your hardware (computer, power supply, internet connection) is working during the exam&#039;&#039;. There will not be extra time granted on the basis of individual connection problems.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Your computer must be able to connect to DTU wi-fi. &amp;lt;!-- &#039;&#039;&#039;Use the standard DTU net which will be open during your exam; NOT the special exam net.&#039;&#039;&#039; --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Software ==&lt;br /&gt;
; A modern Internet Browser ([http://www.mozilla.com/ FireFox], Edge, Safari (Mac only), [http://www.google.com/chrome Google Chrome], [http://www.opera.com/ Opera]).&lt;br /&gt;
: NB: You &#039;&#039;must&#039;&#039; have FireFox, Chrome, or Opera on your machine, so you are able to switch browser in case Edge (Windows) or Safari (Mac) has trouble with a certain website.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Java&lt;br /&gt;
: Download and install from https://www.oracle.com/java/technologies/downloads/ (free) or https://adoptium.net/ (free), in case it is not already on your system. &lt;br /&gt;
: Test that it works:&lt;br /&gt;
:* jEdit, Jalview, and FigTree (see below) must be able to run.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
; Geany (or another GOOD plain text editor)&lt;br /&gt;
: Download and install from http://geany.org/, see the [[Plain text files and Geany]] exercise.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
: If you have unsolvable problems with jEdit, a good alternative is [ Geany].&lt;br /&gt;
: You MUST be able to handle text files with Unix, Mac, or Windows line endings — therefore, Notepad under Windows is NOT good enough.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; PyMol&lt;br /&gt;
: Download and install from https://pymol.org/, see the [[Media:PyMOL_tutorial.pdf|PyMol tutorial]] and the exercises in [[Protein Structure and Visualization|PDB &amp;amp; PyMol]], and [[Exercise:Malaria Vaccine|Malaria vaccine]]. &amp;lt;!-- , and [[ExPSIBLAST|PSI-BLAST]]. --&amp;gt;&lt;br /&gt;
: A license file is found on Learn → Content → Week 06. &lt;br /&gt;
&lt;br /&gt;
; Seaview&lt;br /&gt;
: Download and install from http://doua.prabi.fr/software/seaview, see [[Exercise: Multiple Alignments (Seaview version)]].&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Jalview&lt;br /&gt;
: Download and install from https://www.jalview.org/, see the exercise in [[Exercise: Multiple Alignments (English version)|Multiple Alignments]]. &lt;br /&gt;
; FigTree&lt;br /&gt;
: Download and install from https://github.com/rambaut/figtree/releases, see [[Exercise: Phylogeny]].&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; Text processing software&lt;br /&gt;
: for writing your answers. You can e.g. use:&lt;br /&gt;
:* Microsoft Word &lt;br /&gt;
:* [http://www.openoffice.org/ OpenOffice] / [http://www.neooffice.org/ NeoOffice] / [http://www.libreoffice.org/ LibreOffice] &lt;br /&gt;
:* Pages (for Mac).&lt;br /&gt;
:* [https://docs.google.com/ Google Docs] &lt;br /&gt;
&lt;br /&gt;
; Tool for making PDF files&lt;br /&gt;
: included in Windows 10/11 and Mac. &lt;br /&gt;
&lt;br /&gt;
; Tool for taking screenshots&lt;br /&gt;
: included in Microsoft Word, Windows 10/11, and Mac. &lt;br /&gt;
: For Windows users we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=848</id>
		<title>Checklist for computers</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Checklist_for_computers&amp;diff=848"/>
		<updated>2026-08-20T09:28:29Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Internet connection */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;At the exam in Introduction to Bioinformatics, you are going to use the same resources (web-servers and programs) that you used for the exercises — if your computer has worked fine for all the exercises, it will also work fine for the exam.&lt;br /&gt;
&lt;br /&gt;
== Hardware ==&lt;br /&gt;
; A laptop.&lt;br /&gt;
: It makes no difference whether you use Windows, Mac, or Linux, as long as you have the listed software. However, an iPad, an Android tablet, or a Chromebook will NOT be enough for the exam.&lt;br /&gt;
; A mouse.&lt;br /&gt;
: A mouse is important in order to use PyMOL (see below) optimally. The mouse should have two buttons plus a scroll-wheel in the middle.&lt;br /&gt;
&lt;br /&gt;
== Internet connection ==&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
&#039;&#039;&#039;NOTE 2021:&#039;&#039;&#039; Due to the corona situation, your exam will be from home. Whether you use wireless or cabled connection is not important, but &#039;&#039;it is your own responsibility that your hardware (computer, power supply, internet connection) is working during the exam&#039;&#039;. There will not be extra time granted on the basis of individual connection problems.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Your computer must be able to connect to DTU wi-fi. &amp;lt;!-- &#039;&#039;&#039;Use the standard DTU net which will be open during your exam; NOT the special exam net.&#039;&#039;&#039; --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Software ==&lt;br /&gt;
; A modern Internet Browser ([http://www.mozilla.com/ FireFox], Edge, Safari (Mac only), [http://www.google.com/chrome Google Chrome], [http://www.opera.com/ Opera]).&lt;br /&gt;
: NB: You &#039;&#039;must&#039;&#039; have FireFox, Chrome, or Opera on your machine, so you are able to switch browser in case Edge (Windows) or Safari (Mac) has trouble with a certain website.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Java&lt;br /&gt;
: Download and install from https://www.oracle.com/java/technologies/downloads/ (free) or https://adoptium.net/ (free), in case it is not already on your system. &lt;br /&gt;
: Test that it works:&lt;br /&gt;
:* jEdit, Jalview, and FigTree (see below) must be able to run.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
; Geany (or another GOOD plain text editor)&lt;br /&gt;
: Download and install from http://geany.org/, see the [[Plain text files and Geany]] exercise.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
: If you have unsolvable problems with jEdit, a good alternative is [ Geany].&lt;br /&gt;
: You MUST be able to handle text files with Unix, Mac, or Windows line endings — therefore, Notepad under Windows is NOT good enough.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; PyMol&lt;br /&gt;
: Download and install from https://pymol.org/, see the [https://teaching.healthtech.dtu.dk/material/22111/PyMol_tutorial2017_v4.pdf PyMol tutorial] and the exercises in [[Protein Structure and Visualization|PDB &amp;amp; PyMol]], and [[Exercise:Malaria Vaccine|Malaria vaccine]]. &amp;lt;!-- , and [[ExPSIBLAST|PSI-BLAST]]. --&amp;gt;&lt;br /&gt;
: A license file is found on Learn → Content → Week 06. &lt;br /&gt;
&lt;br /&gt;
; Seaview&lt;br /&gt;
: Download and install from http://doua.prabi.fr/software/seaview, see [[Exercise: Multiple Alignments (Seaview version)]].&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
; Jalview&lt;br /&gt;
: Download and install from https://www.jalview.org/, see the exercise in [[Exercise: Multiple Alignments (English version)|Multiple Alignments]]. &lt;br /&gt;
; FigTree&lt;br /&gt;
: Download and install from https://github.com/rambaut/figtree/releases, see [[Exercise: Phylogeny]].&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
; Text processing software&lt;br /&gt;
: for writing your answers. You can e.g. use:&lt;br /&gt;
:* Microsoft Word &lt;br /&gt;
:* [http://www.openoffice.org/ OpenOffice] / [http://www.neooffice.org/ NeoOffice] / [http://www.libreoffice.org/ LibreOffice] &lt;br /&gt;
:* Pages (for Mac).&lt;br /&gt;
:* [https://docs.google.com/ Google Docs] &lt;br /&gt;
&lt;br /&gt;
; Tool for making PDF files&lt;br /&gt;
: included in Windows 10/11 and Mac. &lt;br /&gt;
&lt;br /&gt;
; Tool for taking screenshots&lt;br /&gt;
: included in Microsoft Word, Windows 10/11, and Mac. &lt;br /&gt;
: For Windows users we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots.&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=847</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=847"/>
		<updated>2026-08-17T17:48:02Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Dec 15 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
Here you can find a guide to the Digital Exam interface (in Danish and English): https://student.dtu.dk/en/exam/exam-guides&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=846</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=846"/>
		<updated>2026-08-17T17:34:50Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Dec 16 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 15 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 15.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=845</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=845"/>
		<updated>2026-08-17T16:12:54Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Dec 16 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=844</id>
		<title>22111 - Introduction to Bioinformatics</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=844"/>
		<updated>2026-08-17T13:08:34Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Course programme */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Formerly known as 36611 and 27611&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Practical information ==&lt;br /&gt;
DTU&#039;s Studies Handbook about [http://www.kurser.dtu.dk/course/22111 #22111]&lt;br /&gt;
&lt;br /&gt;
This course is &#039;&#039;&#039;taught in English since 2019&#039;&#039;&#039; (previously taught in Danish), and it is a practically oriented, introductory 3rd semester (previously 4th semester) course. All students from DTU and other universities are welcome.&lt;br /&gt;
&lt;br /&gt;
For more information, please contact: Associate Professor [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] ([mailto:henni@dtu.dk henni@dtu.dk]), Associate Professor [https://www.dtu.dk/Person/cwis?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] ([mailto:carolet@dtu.dk carolet@dtu.dk]), or Study Coordinator [https://www.dtu.dk/person/line-juul-larsen?id=218401&amp;amp;entity=profile Line Juul Larsen] ([mailto:ljula@dtu.dk ljula@dtu.dk]).&lt;br /&gt;
&lt;br /&gt;
If you want to participate in the course, please sign up through the Studies Division (&amp;quot;studiekontoret&amp;quot;) at DTU.&lt;br /&gt;
If you are not enrolled at the Technical University of Denmark, you have to sign up as a guest student (more information here: [https://www.dtu.dk/english/education/incoming-students/guests Enroll as a guest student at DTU])&lt;br /&gt;
&lt;br /&gt;
== Course programme ==&lt;br /&gt;
&#039;&#039;&#039;Current:&#039;&#039;&#039; &lt;br /&gt;
* [[22111: Course plan autumn 2026]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Previous:&#039;&#039;&#039;&lt;br /&gt;
* [[22111: Course plan autumn 2025]]&lt;br /&gt;
* [[22111: Course plan autumn 2024]]&lt;br /&gt;
* [[22111: Course plan autumn 2023]]&lt;br /&gt;
* [[22111: Course plan autumn 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2021]]&lt;br /&gt;
* [[22111: Course plan spring 2020]]&lt;br /&gt;
* [[22111: Course plan spring 2019]]&lt;br /&gt;
* [[22111:Kursusplan for forår 2018]] (Course Programme Spring 2018, in Danish)&lt;br /&gt;
* [[27611: Kursusplan for forår 2017]] (Course Programme Spring 2017, in Danish)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Special editions:&#039;&#039;&#039;&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2022]]&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2024]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=843</id>
		<title>22111 - Introduction to Bioinformatics</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=843"/>
		<updated>2026-08-17T13:06:45Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Practical information */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Formerly known as 36611 and 27611&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Practical information ==&lt;br /&gt;
DTU&#039;s Studies Handbook about [http://www.kurser.dtu.dk/course/22111 #22111]&lt;br /&gt;
&lt;br /&gt;
This course is &#039;&#039;&#039;taught in English since 2019&#039;&#039;&#039; (previously taught in Danish), and it is a practically oriented, introductory 3rd semester (previously 4th semester) course. All students from DTU and other universities are welcome.&lt;br /&gt;
&lt;br /&gt;
For more information, please contact: Associate Professor [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] ([mailto:henni@dtu.dk henni@dtu.dk]), Associate Professor [https://www.dtu.dk/Person/cwis?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] ([mailto:carolet@dtu.dk carolet@dtu.dk]), or Study Coordinator [https://www.dtu.dk/person/line-juul-larsen?id=218401&amp;amp;entity=profile Line Juul Larsen] ([mailto:ljula@dtu.dk ljula@dtu.dk]).&lt;br /&gt;
&lt;br /&gt;
If you want to participate in the course, please sign up through the Studies Division (&amp;quot;studiekontoret&amp;quot;) at DTU.&lt;br /&gt;
If you are not enrolled at the Technical University of Denmark, you have to sign up as a guest student (more information here: [https://www.dtu.dk/english/education/incoming-students/guests Enroll as a guest student at DTU])&lt;br /&gt;
&lt;br /&gt;
== Course programme ==&lt;br /&gt;
&#039;&#039;&#039;Current:&#039;&#039;&#039; &lt;br /&gt;
* [[22111: Course plan autumn 2025]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Previous:&#039;&#039;&#039;&lt;br /&gt;
* [[22111: Course plan autumn 2024]]&lt;br /&gt;
* [[22111: Course plan autumn 2023]]&lt;br /&gt;
* [[22111: Course plan autumn 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2021]]&lt;br /&gt;
* [[22111: Course plan spring 2020]]&lt;br /&gt;
* [[22111: Course plan spring 2019]]&lt;br /&gt;
* [[22111:Kursusplan for forår 2018]] (Course Programme Spring 2018, in Danish)&lt;br /&gt;
* [[27611: Kursusplan for forår 2017]] (Course Programme Spring 2017, in Danish)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Special editions:&#039;&#039;&#039;&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2022]]&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2024]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=842</id>
		<title>22111 - Introduction to Bioinformatics</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=842"/>
		<updated>2026-08-17T13:06:11Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Practical information */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Formerly known as 36611 and 27611&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Practical information ==&lt;br /&gt;
DTU&#039;s Studies Handbook about [http://www.kurser.dtu.dk/course/22111 #22111]&lt;br /&gt;
&lt;br /&gt;
This course is &#039;&#039;&#039;taught in English since 2019&#039;&#039;&#039; (previously taught in Danish), and it is a practically oriented, introductory 3rd semester (previously 4th semester) course. All students from DTU and other universities are welcome.&lt;br /&gt;
&lt;br /&gt;
For more information, please contact: Associate Professor [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] ([mailto:henni@dtu.dk henni@dtu.dk]), Associate Professor [https://www.dtu.dk/Person/cwis?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] ([mailto:carolet@dtu.dk carolet@dtu.dk]), or Study Coordinator [https://www.dtu.dk/person/line-juul-larsen?id=218401&amp;amp;entity=profile Line Juul Larsen] ([mailto:ljula@dtu.dk ljula@dtu.dk]).&lt;br /&gt;
&lt;br /&gt;
If you want to participate in the course, please sign up through the Studies Division (&amp;quot;studiekontoret&amp;quot;) at DTU.&lt;br /&gt;
If you are not enrolled at the Technical University of Denmark, you have to sign up as a guest student (more information here: [https://www.dtu.dk/english/education/incoming-students/guests General Enroll as a guest student at DTU])&lt;br /&gt;
&lt;br /&gt;
== Course programme ==&lt;br /&gt;
&#039;&#039;&#039;Current:&#039;&#039;&#039; &lt;br /&gt;
* [[22111: Course plan autumn 2025]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Previous:&#039;&#039;&#039;&lt;br /&gt;
* [[22111: Course plan autumn 2024]]&lt;br /&gt;
* [[22111: Course plan autumn 2023]]&lt;br /&gt;
* [[22111: Course plan autumn 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2021]]&lt;br /&gt;
* [[22111: Course plan spring 2020]]&lt;br /&gt;
* [[22111: Course plan spring 2019]]&lt;br /&gt;
* [[22111:Kursusplan for forår 2018]] (Course Programme Spring 2018, in Danish)&lt;br /&gt;
* [[27611: Kursusplan for forår 2017]] (Course Programme Spring 2017, in Danish)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Special editions:&#039;&#039;&#039;&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2022]]&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2024]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=841</id>
		<title>22111 - Introduction to Bioinformatics</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111_-_Introduction_to_Bioinformatics&amp;diff=841"/>
		<updated>2026-08-17T12:55:30Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Practical information */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Formerly known as 36611 and 27611&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Practical information ==&lt;br /&gt;
DTU&#039;s Studies Handbook about [http://www.kurser.dtu.dk/course/22111 #22111]&lt;br /&gt;
&lt;br /&gt;
This course is &#039;&#039;&#039;taught in English since 2019&#039;&#039;&#039; (previously taught in Danish), and it is a practically oriented, introductory 3rd semester (previously 4th semester) course. All students from DTU and other universities are welcome.&lt;br /&gt;
&lt;br /&gt;
For more information, please contact: Associate Professor [http://www.dtu.dk/service/telefonbog/person?id=25617&amp;amp;cpid=214126&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Henrik Nielsen] ([mailto:henni@dtu.dk henni@dtu.dk]), Associate Professor [https://www.dtu.dk/Person/cwis?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] ([mailto:carolet@dtu.dk carolet@dtu.dk]), or Course Coordinator [https://www.dtu.dk/service/telefonbog/person?id=136144&amp;amp;cpid=246063&amp;amp;tab=0 Antonia Celinah Majlund Bjørstorp] ([mailto:acmb@dtu.dk acmb@dtu.dk]).&lt;br /&gt;
&lt;br /&gt;
If you want to participate in the course, please sign up through the Studies Division (&amp;quot;studiekontoret&amp;quot;) at DTU.&lt;br /&gt;
If you are not enrolled at the Technical University of Denmark, you have to sign up as a guest student (more information here: [http://www.dtu.dk/Uddannelse/Gaestestuderende.aspx General Information for guest students from other Danish Universities])&lt;br /&gt;
&lt;br /&gt;
== Course programme ==&lt;br /&gt;
&#039;&#039;&#039;Current:&#039;&#039;&#039; &lt;br /&gt;
* [[22111: Course plan autumn 2025]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Previous:&#039;&#039;&#039;&lt;br /&gt;
* [[22111: Course plan autumn 2024]]&lt;br /&gt;
* [[22111: Course plan autumn 2023]]&lt;br /&gt;
* [[22111: Course plan autumn 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2022]]&lt;br /&gt;
* [[22111: Course plan spring 2021]]&lt;br /&gt;
* [[22111: Course plan spring 2020]]&lt;br /&gt;
* [[22111: Course plan spring 2019]]&lt;br /&gt;
* [[22111:Kursusplan for forår 2018]] (Course Programme Spring 2018, in Danish)&lt;br /&gt;
* [[27611: Kursusplan for forår 2017]] (Course Programme Spring 2017, in Danish)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Special editions:&#039;&#039;&#039;&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2022]]&lt;br /&gt;
* [[Bioinformatics in practice, Faroe Islands 2024]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=840</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=840"/>
		<updated>2026-08-17T11:02:22Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Oct 20 — Case: Malaria vaccine */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=839</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=839"/>
		<updated>2026-08-17T11:02:06Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Oct 6 — Sequence information &amp;amp; logo-plots */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=838</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=838"/>
		<updated>2026-08-17T10:49:31Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[[Protein Structure]] exercise&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=837</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=837"/>
		<updated>2026-08-17T09:41:33Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Sep 1 — Introduction, taxonomy and GenBank */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy, and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[https://teaching.healthtech.dtu.dk/22111/index.php/Protein_Structure Protein Structure exercise]&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=836</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=836"/>
		<updated>2026-08-17T09:16:08Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Sep 1 — Introduction, taxonomy and GenBank */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae967 NCBI Taxonomy: enhanced access via NCBI Datasets]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[https://teaching.healthtech.dtu.dk/22111/index.php/Protein_Structure Protein Structure exercise]&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=835</id>
		<title>22111:Course plan autumn 2026</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=22111:Course_plan_autumn_2026&amp;diff=835"/>
		<updated>2026-08-17T09:05:34Z</updated>

		<summary type="html">&lt;p&gt;Henni: /* Tuesday Nov 24 — Neural Networks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== General information ==&lt;br /&gt;
&lt;br /&gt;
=== Where and when ===&lt;br /&gt;
Lectures plus subsequent exercises will take place every Tuesday afternoon during the semester, starting &#039;&#039;&#039;Tuesday Sep 1 at 13:00&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Lectures will be from 13:00 to approx. 14 in &#039;&#039;&#039;TBA&amp;lt;!--Aud. 54, building 208--&amp;gt;&#039;&#039;&#039;, and the exercises will then take place in &#039;&#039;&#039;TBA&amp;lt;!--the group rooms ALC1 (001), ALC2 (011), ALC4 (012), and &amp;quot;Touch Down&amp;quot; (024) also in building 208--&amp;gt;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Teachers ===&lt;br /&gt;
&lt;br /&gt;
* [https://www.dtu.dk/person/henrik-nielsen?id=25617&amp;amp;entity=profile Henrik Nielsen] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/carolina-barra-quaglia?id=142840&amp;amp;entity=profile Carolina Barra Quaglia] &amp;amp;mdash; Associate professor, course responsible.&lt;br /&gt;
* [https://www.dtu.dk/person/rasmus-wernersson?id=18103&amp;amp;entity=profile Rasmus Wernersson] &amp;amp;mdash; Affiliated professor.&lt;br /&gt;
&amp;lt;!-- * [http://www.dtu.dk/service/telefonbog/person?id=5118&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Anders Gorm Pedersen] &amp;amp;mdash; Professor, guest lecturer. Topic: Phylogenetic trees. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Teaching assistants ===&lt;br /&gt;
* [https://www.dtu.dk/person/melanie-randahl-nielsen?id=118686&amp;amp;entity=profile Melanie Randahl Nielsen] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/pilar-ballesteros-cuartero?id=224218&amp;amp;entity=profile Pilar Ballesteros Cuartero] &amp;amp;mdash; PhD student&lt;br /&gt;
* [https://www.dtu.dk/person/mads-vodder-hartmann?id=137701&amp;amp;entity=profile Mads Vodder Hartmann] &amp;amp;mdash; PhD student, stand-in for Melanie in weeks 1+2&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
* [https://www.dtu.dk/person/david-lokjaer-faurdal?id=98246&amp;amp;entity=profile David Lokjær Faurdal] &amp;amp;mdash; PhD student&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Course content ===&lt;br /&gt;
In this course, a large emphasis is placed on the practical usage of bioinformatics databases and tools. A typical lecture will present the theoretical aspects of the topics of the day — sometimes including a small group exercise using pen and paper — and last about an hour. The rest of the time will be spent on practical computer exercises, where the teachers and teaching assistants will be ready to help.&lt;br /&gt;
&lt;br /&gt;
See also [https://kurser.dtu.dk/course/2026-2027/22111 the course base about 22111].&lt;br /&gt;
&lt;br /&gt;
=== Curriculum ===&lt;br /&gt;
There is no formal textbook. The curriculum consists of the exercise guides, supplemented with various papers and chapters which will be made available on this homepage or on DTU Learn. Please note that &#039;&#039;all&#039;&#039; exercise guides are mandatory curriculum — including the &#039;&#039;answers&#039;&#039; to the exercises which will be made available on DTU Learn after each exercise.&lt;br /&gt;
&lt;br /&gt;
=== Computers ===&lt;br /&gt;
====Hardware====&lt;br /&gt;
&#039;&#039;&#039;You must bring your own laptop&#039;&#039;&#039; to the exercises, and it must be able to connect to DTU&#039;s wireless network. The type of computer / operating system is not important; Windows, Mac or Linux will all work fine. An iPad or an Android tablet, on the other hand, will not be good enough. A Chromebook will also not be enough (unless you have succeeded in installing a Linux distribution on it, but in that case we assume you know what you&#039;re doing). &lt;br /&gt;
&lt;br /&gt;
In some of the exercises (&amp;quot;PDB/PyMOL&amp;quot;, &amp;quot;Malaria vaccine&amp;quot;, and &amp;quot;Mock exam&amp;quot;), you will work with the molecular visualization program PyMOL. This is rather difficult to control by a touchpad, so please remember to &#039;&#039;&#039;bring a mouse&#039;&#039;&#039;. The mouse should have two buttons plus a scroll-wheel.&lt;br /&gt;
&lt;br /&gt;
====Software====&lt;br /&gt;
# Most importantly: an updated &#039;&#039;&#039;internet browser&#039;&#039;&#039; (e.g. [http://www.google.com/chrome Google Chrome], [http://www.mozilla.com/ FireFox], [http://www.opera.com/ Opera], [https://www.microsoft.com/edge Edge], or Safari for Mac only). &#039;&#039;&#039;NB:&#039;&#039;&#039; You must have more than one browser installed; Safari for Mac or Edge for Windows may have glitches with some bioinformatics websites, and in those cases it is important to be able to switch to an alternative browser.&lt;br /&gt;
# A plain text editor for working with, e.g., sequence files. We recommend &#039;&#039;&#039;Geany&#039;&#039;&#039;, which you can download for free from https://geany.org/. You will find some tips and installation instructions in the former exercise &amp;quot;[[Plain_text_files_and_Geany]]&amp;quot;.&lt;br /&gt;
Other software will be installed during the exercises.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Note:&#039;&#039;&#039; Previously in the course, we have used some java-based software; but it is our experience that new Macs (with M-series CPUs, a.k.a. ARM chips) often have problems with java. Therefore, we have replaced these programs with other options: &lt;br /&gt;
* [https://geany.org/ Geany] has replaced [http://jedit.org/ jEdit], see the former exercise in [[Plain_text_files_and_Geany|plain text files]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] has replaced [https://www.jalview.org/ Jalview], see the exercise in [[Exercise:_Multiple_Alignments_(Seaview_version)|multiple alignments]].&lt;br /&gt;
* [https://doua.prabi.fr/software/seaview SeaView] &amp;lt;!--(and to some degree the website [https://itol.embl.de/ iTOL])--&amp;gt; has also replaced the software [https://github.com/rambaut/figtree/releases FigTree], see the exercise in [[Exercise: Phylogeny|phylogenetic trees]].&lt;br /&gt;
Be aware that if you are working on old exam sets, they may refer to the old software.&lt;br /&gt;
&lt;br /&gt;
=== Hand-ins ===&lt;br /&gt;
As preparation for the computer-based exam, each participant or group must write a &amp;quot;&#039;&#039;&#039;logbook&#039;&#039;&#039;&amp;quot; with answers to the questions posed in the exercise guides. After the exercise, you should upload the logbook to DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; All hand-ins are per definition group hand-ins. If you work alone, you must form a &amp;quot;group&amp;quot; of one person. &lt;br /&gt;
&amp;lt;!-- It is possible to hand in as a group. We would &#039;&#039;much&#039;&#039; rather receive one group hand-in than a number of identical logbooks. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You decide which software you prefer for writing the logbook — e.g. Microsoft Word, [http://www.libreoffice.org/ LibreOffice] (free), [http://www.openoffice.org/ Apache OpenOffice] (free), Pages for Mac, [https://docs.google.com/ Google Docs] or similar. You should be able to insert &#039;&#039;&#039;screenshots&#039;&#039;&#039; in the logbooks for documentation purposes. Microsoft Word has a built-in screenshot tool. Both Windows 10/11 and Mac OS also have dedicated screenshot tools.&lt;br /&gt;
&amp;lt;!-- For Windows users, however, we recommend the free program [http://getgreenshot.org/ Greenshot] which can not only take screenshots and copy them to the clipboard, but also make simple edits and annotations in the screenshots. --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Regardless of your choice of writing software, the result &#039;&#039;&#039;must be handed in as a PDF file&#039;&#039;&#039;. LibreOffice and Google Docs can make PDFs directly. MacOS and Windows 10/11 have built-in functions for converting any printable file to PDF. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Please do &#039;&#039;not&#039;&#039; copy the questions&#039;&#039;&#039; from the exercise guide to your logbook. The hand-in module on DTU Learn has a system for plagiarism detection, which will raise an alarm if significant portions of your hand-in are identical to documents found on the internet — and that includes the exercise guides.&lt;br /&gt;
&lt;br /&gt;
In case you don&#039;t finish the exercises Tuesday afternoon, there is still a chance to hand in — the deadline for handing in at Learn is &#039;&#039;&#039;Thursday at 13:00&#039;&#039;&#039; each week. The &#039;&#039;&#039;answers to the exercises&#039;&#039;&#039; will become visible at the same moment (Thursday at 13:00). You should read the answers carefully and compare with your own answers.&lt;br /&gt;
&lt;br /&gt;
We do not offer individual feedback on the hand-ins, but we will give a collective feedback before the lecture the next Tuesday, where we address any common mistakes there may have been.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NB:&#039;&#039;&#039; &#039;&#039;The hand-ins do not affect your grade&#039;&#039; — they are mainly meant as a preparation for the exam. They are also a means for us to check the understanding of the teaching; if we can see that many participants have made the same mistake, we will try to explain the issue better at the beginning of the next lecture.&lt;br /&gt;
&lt;br /&gt;
=== Exam ===&lt;br /&gt;
The 22111 exam is electronic; i.e. you must bring your own computer, and you will &#039;&#039;not&#039;&#039; get a paper copy of the questions. &lt;br /&gt;
&lt;br /&gt;
This year, the exam will be Multiple Choice and &#039;&#039;without&#039;&#039; access to the internet.&lt;br /&gt;
&amp;lt;!-- The questions will be made available as a PDF file on the DTU online exam system. &#039;&#039;&#039;The only accepted hand-in format is PDF&#039;&#039;&#039;. --&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
All aids are allowed at the exam; you can bring any books, papers or notes. You will have &#039;&#039;&#039;open access to the internet&#039;&#039;&#039; which includes all the materials and websites we have used during the course. You are also allowed to search information on Google, Wikipedia, etc., but you are &#039;&#039;not&#039;&#039; allowed to communicate with others through e-mail, Facebook, chat, or file sharing websites. The internet traffic will be logged during the exam to ensure that these restrictions are kept.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
Just like in the weekly hand-ins, we kindly ask you: &#039;&#039;Please don&#039;t copy the questions in your answer document&#039;&#039; — that might result in the answer being flagged as plagiarism.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== DTU Learn &amp;amp; Inside ===&lt;br /&gt;
* Link to this year&#039;s DTU Learn page: https://learn.inside.dtu.dk/d2l/home/325121&lt;br /&gt;
* Link to this year&#039;s Campusnet group: https://campusnet.dtu.dk/cnnet/element/875925&lt;br /&gt;
&lt;br /&gt;
=== Evaluation and feedback ===&lt;br /&gt;
We will be very happy to receive comments, suggestions, criticisms, or praise at any time during the semester. You can:&lt;br /&gt;
* send them by email to the teachers, or &lt;br /&gt;
* write them under &amp;quot;General feedback&amp;quot; in &amp;quot;Discussion&amp;quot; in DTU Learn.&lt;br /&gt;
If somebody writes a message in &amp;quot;Discussion&amp;quot;, you can comment on it. If you see a message you agree on, please comment &amp;quot;Agree!&amp;quot; so that we can see that it is not just one person&#039;s opinion. &lt;br /&gt;
&lt;br /&gt;
In addition, we will conduct a mid-term evaluation in [https://evaluering.dtu.dk/ DTU evaluation].&lt;br /&gt;
&lt;br /&gt;
== Lecture &amp;amp; exercise plan ==&lt;br /&gt;
&lt;br /&gt;
Note: This is a &#039;&#039;preliminary&#039;&#039; plan, changes may occur!&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 1 — Introduction, taxonomy and GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lectures:&#039;&#039;&#039;&lt;br /&gt;
:* &#039;&#039;Introduction to the course, bioinformatics, and computers&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:* &#039;&#039;Test of prior knowledge&#039;&#039; — a Vevox session.&lt;br /&gt;
:* &#039;&#039;Evolution, Taxonomy and DNA as Biological Information&#039;&#039; — Rasmus Wernersson.&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; will be made available on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/PDF/Chapter2_Evolution.pdf Brief Introduction to Evolutionary Theory] — Written by Anders Gorm Pedersen.&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Test of prior knowledge:&#039;&#039;&#039; Go to  https://evaluering.dtu.dk/, click &amp;quot;Test of prior knowledge&amp;quot; under 22111, and fill out the form (it&#039;s anonymous). Spend max. 10 minutes on it. --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039;&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:# [[Plain text files and Geany]] &lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:# [[Using the Taxonomy database]]&lt;br /&gt;
:# [[ExGenbank-new|Using the GenBank database]]&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://teaching.healthtech.dtu.dk/material/22111/ELS_bioinformatics.pdf Bioinformatics]&amp;quot; — Encyclopedia entry from 2009.&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1060 Database resources of the National Center for Biotechnology Information in 2026]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026&lt;br /&gt;
:*&amp;quot;[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start]&amp;quot; (NCBI)&lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkae1114 GenBank 2025 update]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Tuesday Sep 9 — GenBank ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;DNA as Biological Information&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/DNA_SequencingTutorial.pdf DNA sequencing tutorial] — source: IDT Tech Vault&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/HandoutEx_BaseCalling_Simple.pdf &amp;quot;Base-calling&amp;quot; exercise (for printing)] [PDF] / [https://teaching.healthtech.dtu.dk/material/22111/BaseCalling_on_screen_version.pdf &amp;quot;Base-calling&amp;quot; exercise (version for on-screen viewing)] [PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExGenbank-new|Using the GenBank database]] &lt;br /&gt;
:&#039;&#039;&#039;Reference material&#039;&#039;&#039; for the exercise: [https://teaching.healthtech.dtu.dk/material/22111/GenBank+FASTA_handout_revised.pdf GenBank + FASTA format] [PDF] &lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[[File:Phone_34.gif‎]] [http://www.youtube.com/watch?v=YgmoHtLGb5c mRNA splicing] (YouTube).&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[http://www.ncbi.nlm.nih.gov/books/NBK44863/ Entrez Sequences Quick Start] (NCBI)&lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1114 &amp;quot;GenBank 2025 update&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 8 — Translation &amp;amp; UniProt ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein databases&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/VirtualRibosome.pdf Virtual Ribosome] — software article (PDF).&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Exercise: Translation - Virtual Ribosome]] &lt;br /&gt;
:#[[Exercise: The protein database UniProt]] &lt;br /&gt;
:&#039;&#039;&#039;Background material&#039;&#039;&#039; (supposedly known): &lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/protein_handout.pdf Levels of protein structure] [PDF]&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/GeneStructure.pdf Overview of eukaryotic gene structure] (PDF).&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*[https://doi.org/10.1093/nar/gkae1010 &amp;quot;UniProt: the Universal Protein Knowledgebase in 2025&amp;quot;] — article from the annual database issue of Nucleic Acids Research, 2025.&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/uniprotkb_quickguide.pdf &amp;quot;A Quick Guide to UniProtKB&amp;quot;] — nice printable overview.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 15 — Pairwise alignment ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Pairwise alignment&#039;&#039; — Henrik Nielsen.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Page 35-55 in Immunological Bioinformatics (PDF: on DTU Learn → General information and files → Textbook excerpt).&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/New_handout_alignscores.pdf Alignment scores]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPairwiseAlignment|Pairwise alignment]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 22 — BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to BLAST&#039;&#039; — Carolina Barra Quaglia.&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; section 3.2.5 → 3.3 (i.e. pages 47-52) in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise:_BLAST3|BLAST]]&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
::[[File:Phone_34.gif‎]] &#039;&#039;&#039;Videos about BLAST from NCBI:&#039;&#039;&#039; (Video introduction to NCBI&#039;s web interface and Expect Values)  [http://www.youtube.com/playlist?list=PLH-TjWpFfWrtjzMCIvUe-YbrlIeFQlKMq NCBI&#039;s YouTube channel]&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Sep 29 — Protein structure, PDB &amp;amp; PyMOL ===&lt;br /&gt;
:&#039;&#039;&#039;Remember to bring a mouse for this day&#039;s exercise.&#039;&#039;&#039; The mouse should have two buttons and a scroll wheel.&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Protein 3D structure&#039;&#039; — Carolina Barra Quaglia&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://en.wikipedia.org/wiki/Protein_structure Protein Structure (Wikipedia)]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://kurser.dtu.dk/course/22117 22117 Protein Structure and Computational Biology]&lt;br /&gt;
&lt;br /&gt;
:&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://pymol.org/ PyMOL] (choose the newest version)&lt;br /&gt;
::&#039;&#039;&#039;Note:&#039;&#039;&#039; you will need the license file found at DTU Learn under this week&#039;s topic. The license is valid for a limited time. If you need PyMOL for educational purposes later in your studies, you can go to https://pymol.org/edu/index.php and register as a student to get your own license file (and if you don&#039;t receive an email after registering, write to help@schrodinger.com). However, if you need PyMOL to make figures for a scientific publication, you will have to pay for a license.&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:#[[Media:PyMOL_tutorial.pdf|PyMol tutorial]] (PDF) — basic usage of PyMOL.&lt;br /&gt;
:#[https://teaching.healthtech.dtu.dk/22111/index.php/Protein_Structure Protein Structure exercise]&lt;br /&gt;
:&#039;&#039;&#039;Extra material:&#039;&#039;&#039; &lt;br /&gt;
:*&amp;quot;[https://doi.org/10.1093/nar/gkaf1187 RCSB Protein Data Bank: Delivering integrative structures alongside experimental structures and computed structure models]&amp;quot; — article from the annual database issue of Nucleic Acids Research, 2026.&lt;br /&gt;
:*[[PyMOL]] — some tips and tricks.&lt;br /&gt;
:*[https://teaching.healthtech.dtu.dk/material/22111/PDF/PyMOL_structure_navigation.pdf PyMOL basics — a small example] (optional extra exercise)&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 6 — Sequence information &amp;amp; logo-plots ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Sequence information &amp;amp; logo-plots&#039;&#039; — Rasmus Wernersson&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# Pages 68-80 in Immunological Bioinformatics (PDF: on DTU Learn). &lt;br /&gt;
:# Pages 1-9 of &amp;quot;&#039;&#039;Information theory primer&#039;&#039;&amp;quot; ([https://teaching.healthtech.dtu.dk/material/22111/PDF/primer-2.72.pdf PDF])&lt;br /&gt;
:#* Read also the appendix on logarithms (especially log&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;) if needed!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for the lecture: [https://teaching.healthtech.dtu.dk/material/22111/Logo_exercise.pdf How to construct sequence logos] (PDF)&lt;br /&gt;
:[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;Mid-term evaluation:&#039;&#039;&#039; Go to https://evaluering.dtu.dk/ and click &amp;quot;Mid-term evaluation&amp;quot; under 22111 [[Image:Emblem-important_tiny.png‎]]&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExSeqLogos|DNA and Peptide Logos]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;Autumn holiday&#039;&#039;&#039; &lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 20 — Case: Malaria vaccine ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Malaria and vaccines&#039;&#039; — [https://cmp.ku.dk/staff/?pure=en/persons/226923 Thomas Lavstsen], Associate Professor, University of Copenhagen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; [http://www.cdc.gov/dpdx/malaria/ Malaria — Causal Agents / Life Cycle]&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise:Malaria Vaccine|Malaria vaccine]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Oct 27 — Weight matrices ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Introduction to prediction methods, especially Weight Matrices&#039;&#039; — Henrik Nielsen&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; Same as Oct 6!&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
&amp;lt;!--:&#039;&#039;&#039;Handouts&#039;&#039;&#039; for the lecture: --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercises:&#039;&#039;&#039; &lt;br /&gt;
:# [https://teaching.healthtech.dtu.dk/material/22111/Estimationofpseudocounts_new+examples.pdf How to estimate pseudo frequencies]  &#039;&#039;&#039;Note&#039;&#039;&#039;: If you solve this manually, just select a couple of amino acids from the table. But if you solve it programmatically (python, Excel, other...), fill out the entire table.&lt;br /&gt;
:# [[Exercise: Construction of sequence logos and weight matrices|Construction of weight matrices]] &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 3 — PSI-BLAST ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;PSI-BLAST&#039;&#039; — Rasmus Wernersson &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[ExPSIBLAST|PSI-BLAST]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 10 — Multiple alignments ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Multiple alignment&#039;&#039; — Henrik Nielsen &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; RevTrans ([https://www.ncbi.nlm.nih.gov/pmc/articles/PMC169015/ article])&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; [[Exercise: Multiple Alignments (Seaview version)|Multiple Alignments]]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 17 — Phylogenetic trees ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;Phylogenetic Reconstruction: Distance Matrix Methods&#039;&#039; — Henrik Nielsen&lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Extra lecture:&#039;&#039;&#039; &#039;&#039;Bioinformatics and Systems Biology in precision medicine&#039;&#039; — Rasmus Wernersson --&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; &lt;br /&gt;
:# &#039;&#039;Introduction to Tree Building&#039;&#039;, PDF on Learn &amp;lt;!-- XXX WHERE? → Slides etc → Lecture12 --&amp;gt;&lt;br /&gt;
:# &#039;&#039;[http://evolution.berkeley.edu/evolibrary/article/phylogenetics_01 Evolutionary trees]&#039;&#039; (minus the section &amp;quot;How to reconstruct an evolutionary tree&amp;quot;)&lt;br /&gt;
:# &#039;&#039;Understanding Evolutionary Trees&#039;&#039;, [https://teaching.healthtech.dtu.dk/material/22111/PDF/understanding_evo_trees.pdf PDF].&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Handout&#039;&#039;&#039; for lecture: [https://teaching.healthtech.dtu.dk/material/22111/PDF/handout_distance.pdf Reconstructing a distance tree] &lt;br /&gt;
&amp;lt;!-- :&#039;&#039;&#039;Software&#039;&#039;&#039; for installation: [https://github.com/rambaut/figtree/releases FigTree tree-viewer]&lt;br /&gt;
::&#039;&#039;&#039;IMPORTANT NOTE&#039;&#039;&#039; for Windows users: Download the &amp;lt;tt&amp;gt;.zip&amp;lt;/tt&amp;gt; file (FigTree.v1.4.4.zip) and unpack it. Then, go to the &amp;quot;lib&amp;quot; subfolder and double-click the &amp;lt;tt&amp;gt;.jar&amp;lt;/tt&amp;gt; file. The &amp;lt;tt&amp;gt;.exe&amp;lt;/tt&amp;gt; file may not work.&lt;br /&gt;
:&#039;&#039;&#039;TEST&#039;&#039;&#039; of the internal webserver we are going to use during the exercise: Please go to https://services.healthtech.dtu.dk/service.php?TreeHugger and click &amp;quot;View &amp;lt;u&amp;gt;example alignment files&amp;lt;/u&amp;gt;&amp;quot;. Then, copy either the &amp;quot;Sample DNA alignment&amp;quot; or the &amp;quot;Sample peptide dataset&amp;quot; and paste it in the TreeHugger input field. Click &amp;lt;u&amp;gt;Submit query&amp;lt;/u&amp;gt; when instructed by the lecturer.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
:&#039;&#039;&#039;Exercise: [[Exercise: Phylogeny (Seaview version)|Phylogeny]]&#039;&#039;&#039; &lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course:&#039;&#039;&#039; &lt;br /&gt;
::* [http://teaching.healthtech.dtu.dk/22115/ 22115 Computational Molecular Evolution]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Nov 24 — Neural Networks ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA&lt;br /&gt;
:&#039;&#039;&#039;Link to advanced course: &#039;&#039;&#039;&lt;br /&gt;
:: [http://teaching.healthtech.dtu.dk/22125/ 22125: Algorithms in bioinformatics]&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 1 — Bioinformatics in practice + mock exam ===&lt;br /&gt;
:&#039;&#039;&#039;Lecture:&#039;&#039;&#039; &#039;&#039;AI, phage discovery and supercomputing&#039;&#039; — [https://globe.ku.dk/staff-list/?pure=en/persons/271131 Bent Petersen, KU]. &lt;br /&gt;
:&#039;&#039;&#039;Curriculum:&#039;&#039;&#039; (None - lean back and enjoy)&lt;br /&gt;
:&#039;&#039;&#039;Slides:&#039;&#039;&#039; on DTU Learn.&lt;br /&gt;
:&#039;&#039;&#039;Exercise:&#039;&#039;&#039; TBA.&lt;br /&gt;
&lt;br /&gt;
== Exam ==&lt;br /&gt;
&lt;br /&gt;
=== Tuesday Dec 16 ===&lt;br /&gt;
&#039;&#039;&#039;Winter exam 2025:&#039;&#039;&#039; Go to https://eksamen.dtu.dk/ and find 22111. &lt;br /&gt;
&lt;br /&gt;
[https://teaching.healthtech.dtu.dk/material/22111/Vejledning-til-digital-eksamen-DE-DK-ENG-revideret-2023-.pdf Here is a guide] to the Digital Exam interface (in Danish and English).&lt;br /&gt;
&lt;br /&gt;
The assignment will be accessible from &#039;&#039;&#039;XX:00&#039;&#039;&#039; on Dec 16.&lt;br /&gt;
&lt;br /&gt;
=== Checklist for computers ===&lt;br /&gt;
Check here whether your computer has all the software needed for the exam: [[Checklist for computers]]&lt;br /&gt;
&lt;br /&gt;
=== Link collection ===&lt;br /&gt;
A quick overview of the websites we have used in the course: [[Link collection]]&lt;br /&gt;
&lt;br /&gt;
=== FAQ ===&lt;br /&gt;
Questions we have received and answered: [[FAQ]]&lt;/div&gt;</summary>
		<author><name>Henni</name></author>
	</entry>
</feed>