<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://teaching.healthtech.dtu.dk/22111/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Raz</id>
	<title>22111 - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://teaching.healthtech.dtu.dk/22111/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Raz"/>
	<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php/Special:Contributions/Raz"/>
	<updated>2026-10-11T05:49:54Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Logo_exercise_2026.pdf&amp;diff=1161</id>
		<title>File:Logo exercise 2026.pdf</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Logo_exercise_2026.pdf&amp;diff=1161"/>
		<updated>2026-10-06T06:56:32Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1160</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1160"/>
		<updated>2026-10-06T06:23:29Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the peptide example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
[[image:LogoLab_Sep2026.png|500px|right]]&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically built for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Built by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentation with the visualization.&lt;br /&gt;
* There is a stand-alone downloadable build available for running locally on your own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peek inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
We&#039;ll introduce it briefly here, and get back to it in more details when we get to work with weight matrices.&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observe what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for your plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
&#039;&#039;&#039;2026 NOTE:&#039;&#039;&#039; You can reset the web app to defaults by reloading the page. (We&#039;ll make sure to add a &amp;quot;reset&amp;quot; button later).&lt;br /&gt;
&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them later in the course, but they are essentially a way to have a best guess of amino-acid frequencies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on and off to see (visually) what that do - I&#039;ll return to them later in the course.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the Logo Lab resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in titles etc)&lt;br /&gt;
&amp;lt;!-- *# The y-axis scaling must be the same across the 3 plots --&amp;gt;&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the &amp;quot;Chemistry&amp;quot; coloring means (see below the plot) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using while working with weight matrices) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to later in the course when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1157</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1157"/>
		<updated>2026-10-01T19:51:50Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Part 0: Getting to know &amp;quot;LogoLab&amp;quot; */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
[[image:LogoLab_Sep2026.png|500px|right]]&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically built for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Built by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentation with the visualization.&lt;br /&gt;
* There is a stand-alone downloadable build available for running locally on your own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peek inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
We&#039;ll introduce it briefly here, and get back to it in more details when we get to work with weight matrices.&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observe what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for your plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them later in the course, but they are essentially a way to have a best guess of amino-acid frequencies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on and off to see (visually) what that do - I&#039;ll return to them later in the course.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using while working with weight matrices) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to later in the course when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1151</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1151"/>
		<updated>2026-10-01T06:24:10Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Seq2Logo (CBS/DTU) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
[[image:LogoLab_Sep2026.png|500px|right]]&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
We&#039;ll introduce it briefly here, and get back to it in more details when we get to work with weight matrices.&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using next week) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1150</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1150"/>
		<updated>2026-10-01T06:22:53Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Seq2Logo (CBS/DTU) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
[[image:LogoLab_Sep2026.png|500px|right]]&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using next week) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1149</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1149"/>
		<updated>2026-10-01T06:22:44Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* LogoLab (DTU) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
[[image:LogoLab_Sep2026.png|500px|right]]&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using next week) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:LogoLab_Sep2026.png&amp;diff=1148</id>
		<title>File:LogoLab Sep2026.png</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:LogoLab_Sep2026.png&amp;diff=1148"/>
		<updated>2026-10-01T06:20:51Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1147</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1147"/>
		<updated>2026-09-30T14:56:29Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Introducing Seq2Logo */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;2026 note:&#039;&#039;&#039; The new LogoLab webapp also provides the needed functionality for this part, but it&#039;s a little less tested on KL logos. Once you have done the part below with seq2logo (which we&#039;ll also be using next week) you&#039;re welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1146</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1146"/>
		<updated>2026-09-30T14:53:06Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Signal peptides */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* LogoLab quick link: https://wernersson.dk/logolab/&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from LogoLab?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1145</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1145"/>
		<updated>2026-09-30T14:51:48Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Looking for much weaker DNA motifs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1144</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1144"/>
		<updated>2026-09-30T14:23:35Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Shine-Dalgarno sequence */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG. &lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Alignment range&#039;&#039;&#039; (found below the sequence input text field) to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1143</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1143"/>
		<updated>2026-09-30T14:18:23Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* HUMAN donor sites */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Add a title to you plot using the &amp;quot;display options&amp;quot; and include it your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1142</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1142"/>
		<updated>2026-09-30T14:17:14Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* HUMAN donor sites */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;Numbering starts at&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1141</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1141"/>
		<updated>2026-09-30T14:16:32Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that &#039;&#039;&#039;you can set a title for you plot&#039;&#039;&#039; using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1140</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1140"/>
		<updated>2026-09-30T14:16:15Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that **you can set a title for you plot** using the display options.&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1139</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1139"/>
		<updated>2026-09-30T11:52:03Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* HUMAN donor sites */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1138</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1138"/>
		<updated>2026-09-30T11:49:28Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* HUMAN donor sites */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Use the &#039;&#039;&#039;LogoLab&#039;&#039;&#039; webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (PNG download is an easy way to get it).&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1137</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1137"/>
		<updated>2026-09-30T11:47:40Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Part 1: DNA logos */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos (scientific case) =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1136</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1136"/>
		<updated>2026-09-30T11:46:53Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Part 0: Getting to know &amp;quot;LogoLab&amp;quot; */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
[[image:example_sequence-logo-protein-shannon.png|right|border|400px]]&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Example_sequence-logo-protein-shannon.png&amp;diff=1135</id>
		<title>File:Example sequence-logo-protein-shannon.png</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Example_sequence-logo-protein-shannon.png&amp;diff=1135"/>
		<updated>2026-09-30T11:46:26Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1134</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1134"/>
		<updated>2026-09-30T11:45:22Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|border|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1133</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1133"/>
		<updated>2026-09-30T11:45:12Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|frame|400px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1132</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1132"/>
		<updated>2026-09-30T11:44:52Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|500px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
== Investigating the peptide example ==&lt;br /&gt;
&lt;br /&gt;
Load the peptide example (click the &amp;quot;protein&amp;quot; button), and play around with the logo much like in the DNA example:&lt;br /&gt;
# Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.&lt;br /&gt;
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.&lt;br /&gt;
# With the KL model active, click though some of the positions where many AA&#039;s are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.&lt;br /&gt;
# Pseudo-counts: we&#039;ll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I&#039;ll return to them next week.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1131</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1131"/>
		<updated>2026-09-30T11:32:38Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|right|500px]]&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1130</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1130"/>
		<updated>2026-09-30T11:32:08Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
[[image:Example_sequence-logo-dna-shannon.png|center|800px]]&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1129</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1129"/>
		<updated>2026-09-30T11:31:37Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
[[image:sequence-logo-dna-shannon.png|center|800px]]&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Example_sequence-logo-dna-shannon.png&amp;diff=1128</id>
		<title>File:Example sequence-logo-dna-shannon.png</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Example_sequence-logo-dna-shannon.png&amp;diff=1128"/>
		<updated>2026-09-30T11:30:13Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1127</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1127"/>
		<updated>2026-09-30T11:28:47Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Add error-bars to your plot and, play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1126</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1126"/>
		<updated>2026-09-30T11:28:00Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Investigating the DNA example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
# Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
# Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
# Play around with the &amp;quot;Display option&amp;quot; to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use).&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1125</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1125"/>
		<updated>2026-09-30T11:26:19Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Part 0: Getting to know &amp;quot;LogoLab&amp;quot; */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
Open up LogoLab in a different tab / window: https://wernersson.dk/logolab/&lt;br /&gt;
&lt;br /&gt;
== Investigating the DNA example ==&lt;br /&gt;
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works. &lt;br /&gt;
&lt;br /&gt;
With the DNA example do the following:&lt;br /&gt;
1) Play around with the &#039;&#039;&#039;logo model&#039;&#039;&#039; - make sure you understand what each model is about (cross-ref with the lecture slides).&lt;br /&gt;
2) Return to &amp;quot;Shannon information&amp;quot; as the model, and observer what happens in the &amp;quot;Inspector&amp;quot; view (scroll down a little bit to see it) when you click a position in the LOGO.&lt;br /&gt;
3) Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.&lt;br /&gt;
4)&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1124</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1124"/>
		<updated>2026-09-30T11:14:19Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Part 1: DNA logos */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 0: Getting to know &amp;quot;LogoLab&amp;quot; =&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1123</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1123"/>
		<updated>2026-09-30T11:13:12Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* LogoLab (DTU) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1122</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1122"/>
		<updated>2026-09-30T11:12:49Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* WebLogo (Berkeley) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== LogoLab (DTU) ==&lt;br /&gt;
For this exercise we&#039;ll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally:&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it&#039;ll migrate to a DTU server soon).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Points of notice:&#039;&#039;&#039;&lt;br /&gt;
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.&lt;br /&gt;
* The is a stand-alone downloadable build available for running locally on you own machine.&lt;br /&gt;
* It comes with an &amp;quot;Inspector&amp;quot; functionality that allows you to &amp;quot;peak inside&amp;quot; each position in the LOGO to see the numbers behind the visuals.&lt;br /&gt;
* It has an extensive help-page going over all the equations used.&lt;br /&gt;
&lt;br /&gt;
It you want a good command-line tool as well, we&#039;ll recommend the &amp;quot;weblogo&amp;quot; program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1121</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1121"/>
		<updated>2026-09-30T11:03:11Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: September 2026)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== WebLogo (Berkeley) ==&lt;br /&gt;
[[Image:Weblogologo.png|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; http://weblogo.berkeley.edu/ (we&#039;ll use &#039;&#039;&#039;WebLogo version 2&#039;&#039;&#039; for this exercise)&lt;br /&gt;
&lt;br /&gt;
A good &#039;&#039;&#039;general-purpose logo&#039;&#039;&#039; generator for BOTH &#039;&#039;&#039;DNA&#039;&#039;&#039; and &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1120</id>
		<title>ExSeqLogos v2</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=ExSeqLogos_v2&amp;diff=1120"/>
		<updated>2026-09-30T11:02:47Z</updated>

		<summary type="html">&lt;p&gt;Raz: Original copy-over of the old exercise&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Exercise written by:&#039;&#039;&#039; [http://www.dtu.dk/service/telefonbog/person?id=18103&amp;amp;cpid=214039&amp;amp;tab=2&amp;amp;qt=dtupublicationquery Rasmus Wernersson] (last update: March 2017)&lt;br /&gt;
= Introduction =&lt;br /&gt;
In this exercise we will introduce two methods for generating &#039;&#039;&#039;sequences logos&#039;&#039;&#039;, and we will investigate, how we can extract and compare sequences information from &#039;&#039;&#039;large sets of sequences&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== WebLogo (Berkeley) ==&lt;br /&gt;
[[Image:Weblogologo.png|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; http://weblogo.berkeley.edu/ (we&#039;ll use &#039;&#039;&#039;WebLogo version 2&#039;&#039;&#039; for this exercise)&lt;br /&gt;
&lt;br /&gt;
A good &#039;&#039;&#039;general-purpose logo&#039;&#039;&#039; generator for BOTH &#039;&#039;&#039;DNA&#039;&#039;&#039; and &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
== Seq2Logo (CBS/DTU) == &lt;br /&gt;
[[Image:CBS_logo.gif‎|right]]&lt;br /&gt;
&#039;&#039;&#039;Link:&#039;&#039;&#039; https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
&lt;br /&gt;
A more advanced method for working with &#039;&#039;&#039;peptide&#039;&#039;&#039; sequences.&lt;br /&gt;
&lt;br /&gt;
= Part 1: DNA logos =&lt;br /&gt;
We&#039;ll start out by investigating a small dataset of &#039;&#039;&#039;human splice sites&#039;&#039;&#039; (donor/acceptor)&lt;br /&gt;
&lt;br /&gt;
http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
&lt;br /&gt;
== HUMAN donor sites ==&lt;br /&gt;
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.&lt;br /&gt;
 CAAAACCATTGTGAGTAATC&lt;br /&gt;
 GCCAGAGCAGGTAAAATATC&lt;br /&gt;
 GAACAGTCAGGTCTGTTGCT&lt;br /&gt;
 GAAGGCCCAGGTGAGCATAA&lt;br /&gt;
 TCCTCTACAGGTGGGTACAT&lt;br /&gt;
 GGCGTCCCGCGTAAGTATGG&lt;br /&gt;
 CCTCGTGCAGGTAAGATTAA&lt;br /&gt;
 TGCATGACAGGTGAGTGTTA&lt;br /&gt;
 GAAATGTACAGTAAGTCTCT&lt;br /&gt;
 GGTTCTCTGGGTAAGTAGAG&lt;br /&gt;
 AAATGTACAGGTGAGTACTG&lt;br /&gt;
 ACCTCGCTTGGTACGTGGGA&lt;br /&gt;
 AATCAGACAGGTATAGAAAC&lt;br /&gt;
 AGGACAGAAGGTAATTTTCT&lt;br /&gt;
 AACTATTTGGGTAGGTAGCA&lt;br /&gt;
 GAACTTCCAGGTGTGTGCAG&lt;br /&gt;
 AAACTTGAAGGTATGTTGTT&lt;br /&gt;
 CTGGGATAAGGTAAAAGTAT&lt;br /&gt;
 TTGCACCCAGGTTAGTGGAT&lt;br /&gt;
 ACTTCAATCGGTATGTTTTC&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 1: Our first DNA LOGO&#039;&#039;&#039;&lt;br /&gt;
* Go to the WebLogo server and click &amp;quot;Create&amp;quot; (or follow the direct link: http://weblogo.berkeley.edu/logo.cgi)&lt;br /&gt;
* Paste in the sequences above and hit &amp;quot;&#039;&#039;&#039;create logo&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #1:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the logo in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
* Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?&lt;br /&gt;
* How many bases are from the EXON and how many are from the INTRON?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK 2: prettifying the LOGO&#039;&#039;&#039;&lt;br /&gt;
* Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the &#039;&#039;&#039;numbering scheme&#039;&#039;&#039; to display the GT to start at position &amp;quot;0&amp;quot;.&lt;br /&gt;
* Play around with the &amp;quot;&#039;&#039;&#039;First position number&#039;&#039;&#039;&amp;quot; setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.&lt;br /&gt;
* While we&#039;re at it, put a title on the LOGO plot to indicate that it&#039;s about human donor sites.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #2:&#039;&#039;&#039;&lt;br /&gt;
* Paste in the new logo plot in your report (if the PNG files gives you problems, try generating a PDF instead)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;TASK/QUESTION #3:&#039;&#039;&#039;&lt;br /&gt;
* Finally for good measure we also want to generate a &#039;&#039;&#039;frequency&#039;&#039;&#039; plot for the human donor sites.&lt;br /&gt;
* Find the option to do this, and paste in the resultant LOGO in your report.&lt;br /&gt;
&lt;br /&gt;
== Research task: cross-species comparison ==&lt;br /&gt;
[[Image:Cogs_brain.png|50px]]&lt;br /&gt;
As we have seen above, it&#039;s pretty straightforward to generate DNA logos (&#039;&#039;provided that the data has already been well prepared&#039;&#039;), and it&#039;s now time to perform a real &#039;&#039;&#039;research task&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;INPUT DATA:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:&lt;br /&gt;
* &#039;&#039;Homo sapiens&#039;&#039; (Human)&lt;br /&gt;
* &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly)&lt;br /&gt;
* &#039;&#039;Arabidopsis thaliana&#039;&#039; (Thale Cress)&lt;br /&gt;
* &#039;&#039;Saccharomyces cerevisiea&#039;&#039; (Baker&#039;s / Brewer&#039;s yeast)&lt;br /&gt;
* &#039;&#039;Schizosaccharomyces pombe&#039;&#039; (Fission yeast)&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/LOGO_exercise_donor+acceptor_Q4.zip LOGO_exercise_donor+acceptor_Q4.zip]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
It&#039;s is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.&lt;br /&gt;
&lt;br /&gt;
[[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; &lt;br /&gt;
* Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it&#039;s important that they are easy to tell apart).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #4:&#039;&#039;&#039;&lt;br /&gt;
* Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?&lt;br /&gt;
&lt;br /&gt;
== Looking for much weaker DNA motifs ==&lt;br /&gt;
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.&lt;br /&gt;
&lt;br /&gt;
=== Shine-Dalgarno sequence ===&lt;br /&gt;
In &#039;&#039;&#039;prokaryotes&#039;&#039;&#039; the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Shine-Dalgarno_sequence Shine-Dalgarno sequence]&#039;&#039;&#039;&amp;quot;. The consensus sequence is &#039;&#039;&#039;AGGAGG&#039;&#039;&#039; (DNA: AGGAGG) and it&#039;s located approximately 8bp before the &#039;&#039;&#039;AUG&#039;&#039;&#039; (DNA: ATG).&lt;br /&gt;
&lt;br /&gt;
In order to investigate this in more details, we have prepared a data set of &#039;&#039;&#039;500 &#039;&#039;E. coli&#039;&#039; genes&#039;&#039;&#039; which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/ecoli_500_genes+50bp_upstream.txt ecoli_500_genes+50bp_upstream.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a LOGO of the 500 &#039;&#039;E. coli&#039;&#039; sequences. Since we&#039;re visualizing 100bps it will be easier to see if we stretch out the LOGO a bit: set the dimensions to 50*5cm.&lt;br /&gt;
* As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generate a new LOGO of the &#039;&#039;E. coli&#039;&#039; data set, but set the &#039;&#039;&#039;Logo range&#039;&#039;&#039; to position 25-75 to only show the middle part of the data.&lt;br /&gt;
* Play around with the dimensions to get a good looking LOGO.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #5:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the start codon always &amp;quot;ATG&amp;quot;?&lt;br /&gt;
* Can you see something resembling the Shine-Dalgarno sequence (at which positions)?&lt;br /&gt;
&lt;br /&gt;
In order to get a better &amp;quot;view&amp;quot; of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039;&lt;br /&gt;
* Set the range to position 30-50, and play around with the dimensions to get a good looking plot.&lt;br /&gt;
* Furthermore, set the Y-axis to be capped at 0.5 bits, and put in Y-axis tic marks at 0.1 bit intervals.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION #6:&#039;&#039;&#039;&lt;br /&gt;
* Paste your LOGO into the report.&lt;br /&gt;
* Is the logo (more or less) consistent with the 6 bp consensus sequence?&lt;br /&gt;
&lt;br /&gt;
=== Kozak sequence ===&lt;br /&gt;
[[Image:Cogs_brain.png|50px|left]]&lt;br /&gt;
As the final step in our work with DNA logos, we shall investigate the signal around &#039;&#039;&#039;eukaryotic&#039;&#039;&#039; START codons, and see if we can find the &amp;quot;&#039;&#039;&#039;[http://en.wikipedia.org/wiki/Kozak_consensus_sequence Kozak sequence]&#039;&#039;&#039;&amp;quot; which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).&lt;br /&gt;
It&#039;s is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.&lt;br /&gt;
&lt;br /&gt;
* [[Image:Document-save.png|left|25px]] &#039;&#039;&#039;Data set:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/scer_50bp_5utr+50bp_CDS.seq_only.top500.txt scer_50bp_5utr+50bp_CDS.seq_only.top500.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 7:&#039;&#039;&#039;&lt;br /&gt;
* Can you find any positions (except for the ATG) with &#039;&#039;&#039;information content&#039;&#039;&#039; above 0.2 bits?&lt;br /&gt;
* For any such position what is the approximate &#039;&#039;&#039;frequency&#039;&#039;&#039; of the most common base?&lt;br /&gt;
* Include relevant LOGO plots in your report&lt;br /&gt;
&lt;br /&gt;
= Part 2: Protein logos =&lt;br /&gt;
We&#039;ll start our work with peptide LOGOs with a study of &#039;&#039;&#039;signal peptides&#039;&#039;&#039;, which is a well understood system, and we have prepared data files for this in advance. &lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
From there we&#039;ll go on to see how we can work with entire protein sequences as part of a &#039;&#039;&#039;protein engineering&#039;&#039;&#039; exercise.&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Signal peptides ==&lt;br /&gt;
&lt;br /&gt;
* WebLogo quick link: http://weblogo.berkeley.edu/logo.cgi&lt;br /&gt;
* Seq2Logo quick link: https://services.healthtech.dtu.dk/services/Seq2Logo-2.0/&lt;br /&gt;
Most secreted proteins have an N-terminal [http://en.wikipedia.org/wiki/Signal_peptide signal peptide] that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.&lt;br /&gt;
&lt;br /&gt;
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and [http://en.wikipedia.org/wiki/Gram-negative_bacteria Gram-negative bacteria] (bacteria with two membranes) have shorter signal peptides than [http://en.wikipedia.org/wiki/Gram-positive_bacteria Gram-positive bacteria] (bacteria with one membrane). (Not much is known about Archaeal signal peptides). &lt;br /&gt;
&lt;br /&gt;
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups. &lt;br /&gt;
&lt;br /&gt;
[[Image:Document-save.png|left|25px]] &#039;&#039;&#039;DATA:&#039;&#039;&#039; Zip archive signal peptides for all 3 taxonomical groups. &#039;&#039;&#039;Download link:&#039;&#039;&#039; [https://teaching.healthtech.dtu.dk/material/22111/files/signal_peptides.zip signal_peptides.zip]&lt;br /&gt;
&lt;br /&gt;
   #Seqs #File name&lt;br /&gt;
    3280 EUK.sp.25+5.fasta&lt;br /&gt;
     416 gram+.sp.25+5.fasta&lt;br /&gt;
     846 gram-.sp.25+5.fasta&lt;br /&gt;
&lt;br /&gt;
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been &#039;&#039;&#039;aligned at the cleavage site&#039;&#039;&#039; as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead. &lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEG CDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLA APSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAP DGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNA ENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLA NVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFS VNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASA SSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEA KTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSA HHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMG QTFQY&lt;br /&gt;
 &lt;br /&gt;
 |.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the &amp;quot;^&amp;quot; mark&lt;br /&gt;
&lt;br /&gt;
=== Signal peptide comparison ===&lt;br /&gt;
&#039;&#039;&#039;OVERALL TASK:&#039;&#039;&#039;&lt;br /&gt;
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to &#039;&#039;&#039;compare&#039;&#039;&#039; the motifs in the signal peptides across the three taxonomical groups.&lt;br /&gt;
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but &#039;&#039;&#039;make sure that&#039;&#039;&#039;:&lt;br /&gt;
*# It&#039;s easy to compare the plots (put in labels etc)&lt;br /&gt;
*# The y-axis scaling must be the same across the 3 plots&lt;br /&gt;
*# Plot the &#039;&#039;&#039;signal peptide part&#039;&#039;&#039; of the sequences in &#039;&#039;&#039;negative numbers&#039;&#039;&#039;&lt;br /&gt;
*# Read up on what the coloring means (click the &amp;quot;?&amp;quot; in &amp;quot;color scheme) - this will help you interpret the results.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 8:&#039;&#039;&#039;&lt;br /&gt;
* Put the 3 plots in your report and comment on the following:&lt;br /&gt;
* Which &#039;&#039;&#039;single positions&#039;&#039;&#039; (using the negative numbering scheme) appears to be the most important across all data sets?&lt;br /&gt;
* Are there any &#039;&#039;&#039;regions&#039;&#039;&#039; where a certain class of amino acids appears to be important? What characterizes these amino acids?&lt;br /&gt;
* Any striking differences between the three taxonomical groups?&lt;br /&gt;
&lt;br /&gt;
-------&lt;br /&gt;
=== Introducing Seq2Logo ===&lt;br /&gt;
As the next step we shall investigate how to work with the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we&#039;ll get back to next week when we work with &#039;&#039;&#039;weight matrices&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; create a basic LOGO plot for the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; dataset&lt;br /&gt;
* Set logo type to &amp;quot;&#039;&#039;&#039;Shannon&#039;&#039;&#039;&amp;quot;&lt;br /&gt;
* Put in a title&lt;br /&gt;
* Leave the rest of the &amp;quot;strange&amp;quot; options as default - for reasons we&#039;ll learn later they matter less for a &#039;&#039;&#039;large&#039;&#039;&#039; and &#039;&#039;&#039;well-balanced&#039;&#039;&#039; data set such as this one&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 9:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* Does it (more or less) show the same as the plot from WebLogo?&lt;br /&gt;
&amp;lt;!--&lt;br /&gt;
=== Kullback-Leibler LOGO ===&lt;br /&gt;
The Seq2Logo server has the option to produce a special type of LOGO plot, which shows both over-representation and under-representation of the amino acids. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a Kullback-Leibler LOGO based on the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set (but do otherwise keep the same options as above).&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include the plot in your report&lt;br /&gt;
* By comparing this plot to the one you have just generated for question 9, it should be easy to see, that the under-represented Amino Acids is visualized &#039;&#039;&#039;below&#039;&#039;&#039; the x-axis.&lt;br /&gt;
* Can you find any regions where polar (color = green) amino acids are under-represented?&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Small data sets ===&lt;br /&gt;
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the &#039;&#039;&#039;eukaryote&#039;&#039;&#039; data set:&lt;br /&gt;
&lt;br /&gt;
 --MKQILLISLVVVLAVFAFNVAEGCDATC&lt;br /&gt;
 KTHSVFGFFFKVLLIQVYLFNKTLAAPSPI&lt;br /&gt;
 ---MARGAALALLLFGLLGVLVAAPDGGFD&lt;br /&gt;
 MARMKYNIALIGILASVLLTIAVNAENACN&lt;br /&gt;
 -------MSFRSLLALSGLVCSGLANVISK&lt;br /&gt;
 --MSSTWIKFLFILTLVLLPYSVFSVNIFA&lt;br /&gt;
 ---------MIALFVLMGLMAAASASSCCS&lt;br /&gt;
 KKTVAALSFLFIVLFVAQEIAVTEAKTCEN&lt;br /&gt;
 -----MKAIFVSALLVVALVASTSAHHQEL&lt;br /&gt;
 ------MLRLLLLPLFLFTLSMCMGQTFQY&lt;br /&gt;
 MFRVTSVGCLLLVIVFLNLVVPTSACRAEG&lt;br /&gt;
 GRAMVARLGLGLLLLALLLPTQIYCNQTSV&lt;br /&gt;
 ---MKNHLLFWGVLAVFIKAVHVKAQEDER&lt;br /&gt;
 ----MQRLCVCVLILALALTAFSEASWKPR&lt;br /&gt;
 ---MKLFTTLSASLIFIHSLGSTRAAPVTG&lt;br /&gt;
 LRLLLSALKPGIHVPRAGPAAAFGTSVTSA&lt;br /&gt;
 ---------MKSLIVFACLVAYAAADCTSL&lt;br /&gt;
 ------MKTALPLLLLTCLVAAVQSTGSQG&lt;br /&gt;
 --MGLRALMLWLLAAAGLVRESLQGEFQRK&lt;br /&gt;
 -MATTRFPSLLFYSYIFLLCNGSMAQLFGQ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;TASK:&#039;&#039;&#039; Generate a standard LOGO of the 20 sequences&lt;br /&gt;
* [[Image:Emblem-important_tiny.png‎]] &#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; Be sure to reset the &#039;&#039;&#039;Seq2Logo&#039;&#039;&#039; page, so that you&#039;re not accidentally calculating the plots with the previous input data&lt;br /&gt;
* Set type to &amp;quot;Shannon&amp;quot;&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;0&#039;&#039;&#039; (this will disable the advanced functionality for small sequence sets, and produce a bare-bones LOGO)&lt;br /&gt;
* Generate the LOGO - and &#039;&#039;&#039;save it for comparison&#039;&#039;&#039; (or just paste it into your report right away).&lt;br /&gt;
&lt;br /&gt;
This logo represented the information there can be obtained from this limited set of sequences using the standard LOGO algorithm.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Follow-up TASK:&#039;&#039;&#039; Generate a pseudo-count assisted LOGO of the 20 sequences&lt;br /&gt;
* Set &#039;&#039;&#039;Weight on prior&#039;&#039;&#039; to &#039;&#039;&#039;200&#039;&#039;&#039;&lt;br /&gt;
* Generate the new LOGO and compare it to the one we created a moment ago.&lt;br /&gt;
&lt;br /&gt;
[[Image:Office-notes-line_drawing.png|30px|left]]&lt;br /&gt;
&#039;&#039;&#039;QUESTION 10:&#039;&#039;&#039;&lt;br /&gt;
* Include BOTH plots in your report&lt;br /&gt;
* Comment &#039;&#039;&#039;briefly&#039;&#039;&#039; on what you see - which of the plots best resemble what we learned from the big data sets?&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=969</id>
		<title>Using the Taxonomy database</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=969"/>
		<updated>2026-08-30T19:13:25Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* The NCBI Taxonomy Database */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: &#039;&#039;&#039;Rasmus Wernersson&#039;&#039;&#039; and &#039;&#039;&#039;Henrik Nielsen&#039;&#039;&#039;.&lt;br /&gt;
 &lt;br /&gt;
==Background==&lt;br /&gt;
When comparing DNA and protein sequences from different species it is important to keep in mind that all living organisms at some point in time has shared a common ancestor. Some organisms are closely related and have recently derived from a common ancestor (e.g. Human and Chimpanzee, which diverged 5-10 million years ago) and some are more distantly related (e.g. Human and mouse, diverged 100-150 million years ago).&lt;br /&gt;
&lt;br /&gt;
The more closely related two organisms are, the more similar their sequences will be (say, when comparing the Alpha Globin gene from each of the organisms), and the more likely it will be that similar looking genes from each organism still have the same function (MUCH more about this when we get to pairwise alignment and BLAST searches).&lt;br /&gt;
&lt;br /&gt;
[[file:ApesTax_600.png‎|center|frame|Apes taxonomy (Detailed taxonomy of the Great Apes: Human, Chimp, Gorilla, Orangutan - from Wikipedia)]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Phylogeny vs. taxonomy===&lt;br /&gt;
As we discussed in the lecture all life is organised in a hierarchical taxonomical system, which approximates the &amp;quot;true&amp;quot; underlying phylogeny to a large degree. It&#039;s therefore often important to know where a specific organism is placed in the taxonomical system - this type of information will also always be included along with DNA/Protein sequences from the big databases such as GenBank and UniProt.&lt;br /&gt;
&lt;br /&gt;
Today we will explore various ways to look up and compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
===A word about Wikipedia===&lt;br /&gt;
The free online encyclopedia [http://en.wikipedia.org/wiki/Main_Page Wikipedia] (and other similar resources) is a GREAT way to start out when you need to look up information about a new topic - in this case taxonomy. Almost all species entries in Wikipedia has a &amp;quot;Scientific classification&amp;quot; box which includes taxonomical information (for example see the entry on [http://en.wikipedia.org/wiki/Orangutan Orangutan] or [http://en.wikipedia.org/wiki/Fucus_vesiculosus Fucus vesiculosus] (Bladder wrack / Blæretang)).&lt;br /&gt;
&lt;br /&gt;
HOWEVER: Keep in mind that Wikipedia is NOT a reliable source of information, even if most entries are of a very good quality. The facts in the Wikipedia entries have not been verified by taxonomy experts and can potentially be wrong (everybody can go in and edit the text). We need to look up the taxonomy in an official database (in this case we&#039;ll be using NCBI Taxonomy) before you can state it as a fact.&lt;br /&gt;
&lt;br /&gt;
You CANNOT quote Wikipedia as the only source of your information - you&#039;ll need to find the original primary source of the information or look it up in an official database.&lt;br /&gt;
&lt;br /&gt;
===A word about AI===&lt;br /&gt;
As with Wikipedia using AI to generate a taxonomical analysis (e.g. comparing how a bunch of species are related) can be a great way to get an overview, and a (typically) well written explanation. You will need to ask the AI to include references &#039;&#039;&#039;to actual scientific sources&#039;&#039;&#039; to document the validity of the results, and &#039;&#039;&#039;you will be responsible&#039;&#039;&#039; for double checking the AI output and making sure not only the data is correct, but also that the conclusions are sound.&lt;br /&gt;
&lt;br /&gt;
===The need for a Ground Truth===&lt;br /&gt;
Luckily, modern taxonomy is very well established and there an internationally recognized organization that makes revisions in a highly regulated manner. This &#039;&#039;&#039;taxonomy standard&#039;&#039;&#039; is captured in the &#039;&#039;&#039;NCBI Taxonomy&#039;&#039;&#039; database we&#039;ll be working with in the sections below. For all molecular data NCBI Taxonomy is THE standard reference.&lt;br /&gt;
&lt;br /&gt;
==The NCBI Taxonomy Database==&lt;br /&gt;
[[File:NcbiTax2026.jpg|800px|center|border]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;Main link:&#039;&#039;&#039; https://www.ncbi.nlm.nih.gov/datasets/taxonomy/&lt;br /&gt;
&lt;br /&gt;
(&#039;&#039;2026 note: NCBI is in process of migrating to this new site - if you find the old NCBI Tax homepage via Google (or other means), make sure to press the link they provide to jump to the new site&#039;&#039;).&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As mentioned above NCBI Taxonomy will serve as our &#039;&#039;&#039;Ground Truth&#039;&#039;&#039; for everything related to taxonomy. The NCBI Tax database provides the numerical enumeration of species (and other taxonomical levels) that is used and referenced in most Sequence databases, such as GenBank (DNA) and UniProt (Protein). For example human (&#039;&#039;Homo sapiens&#039;&#039;) has the ID &amp;quot;&#039;&#039;&#039;9606&#039;&#039;&#039;&amp;quot; and Yeast (&#039;&#039;Saccharomyces cerevisiae&#039;&#039;) as the ID &amp;quot;&#039;&#039;&#039;4932&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
NCBI Tax is perhaps not a database you would browse for fun (depending on your level of geekiness). It&#039;s good for looking up definitions, and for comparing the taxonomical position of multiple organisms (since the information is so densely presented).&lt;br /&gt;
&lt;br /&gt;
===Example: Homo sapiens===&lt;br /&gt;
# Open the NCBI Taxonomy webpage in a new browser window/tab (see link above)&lt;br /&gt;
# Search for &amp;quot;&#039;&#039;Homo sapiens&#039;&#039;&amp;quot;.&lt;br /&gt;
## Dont Panic: An enormous amount of information is shown - for example about genome sequences. In this case we only need to look at the information presented in the &amp;quot;&#039;&#039;&#039;Taxonomy&#039;&#039;&#039;&amp;quot; tab.&lt;br /&gt;
## Notice the Taxonomy ID - &#039;&#039;&#039;9606&#039;&#039;&#039; as mentioned above.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Lineage&#039;&#039;&#039;&amp;quot; (box on the right hand side): Here a condensed overview of the human lineage is shown by default. Notice the taxonomical ranks we talked about in the lecture (&amp;quot;Phylum&amp;quot;, &amp;quot;Class&amp;quot; etc). You can navigate to the definition of these groups by clicking on them.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Full lineage&#039;&#039;&#039;&amp;quot;: Click this to view a FULL list of all the groups leading &amp;quot;down&amp;quot; to human. Notice, that you can &amp;quot;mouse over&amp;quot; the groups to see a pop-up with the taxonomical rank. NOTICE: The rank &amp;quot;CLADE&amp;quot; is used whenever a group does not have a common English name (&amp;quot;Clade&amp;quot; is the generic name for a uniquely defined taxonomical group).&lt;br /&gt;
&lt;br /&gt;
Play around with the Homo sapiens page for a bit to familiarize yourself with the interface, and answer the following questions along the way:&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1a&#039;&#039;&#039;: What is the TaxID of &amp;quot;&#039;&#039;Metazoa&#039;&#039;&amp;quot;? &lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1b&#039;&#039;&#039;: What is the family that contains humans?&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1c&#039;&#039;&#039;: Are humans vertebrates? (Latin: Vertebrata)?&lt;br /&gt;
&lt;br /&gt;
===Comparing taxonomy using NCBI Tax===&lt;br /&gt;
[[Image:Fruitfly.jpg|thumb|400px|right|Fruit fly (&#039;&#039;Drosophila melanogaster&#039;&#039;) - source: [http://en.wikipedia.org/wiki/Drosophila_melanogaster Wikipedia] ]]Besides being useful for being the official database behind the TaxID&#039;s used in GenBank (and other databases), NCBI Tax actually makes it easy to compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
Let&#039;s take the situation where you have read an interesting paper comparing a DNA sequence between the following three organisms: &#039;&#039;Homo sapiens&#039;&#039; (Human), &#039;&#039;Mus musculus&#039;&#039; (Mouse), and &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly), but you have no idea about the relationship between the three organisms. &lt;br /&gt;
&lt;br /&gt;
We can look this up in NCBI Tax:&lt;br /&gt;
&lt;br /&gt;
# Open two browser windows/tabs and search for &#039;&#039;&#039;Homo sapiens&#039;&#039;&#039; and &#039;&#039;&#039;Mus musculus&#039;&#039;&#039;.&lt;br /&gt;
# By comparing the &amp;quot;lineage&amp;quot; text it will be easy to find out at which taxonomical level human and mouse differ.&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2A&#039;&#039;&#039;: Look at the (simple/abbreviated) lineage information and find lowest ranking common group for human and mouse - what is the name and what is the rank?&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2B&#039;&#039;&#039;: Look at the &amp;quot;&#039;&#039;&#039;full lineage&#039;&#039;&#039; and find the lowest ranking group human and mouse have in common (it&#039;s OK if the rank is &amp;quot;clade&amp;quot;). What is the name and TaxId of the group?&lt;br /&gt;
&lt;br /&gt;
Now repeat the analysis by comparing Human and Fruit Fly (&#039;&#039;Drosophila melanogaster&#039;&#039;)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3A&#039;&#039;&#039;: What is the lowest rank group shared in the (simplified/abbreviated) lineage? (TaxID, Name)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3B&#039;&#039;&#039;: What is the lowest rank group shared in the full lineage? (TaxID, Name)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Fishing in NCBI Tax using the Common Tree function===&lt;br /&gt;
[[Image:Zebrafisch.jpg|thumb|400px|right|Zebrafish (&#039;&#039;Danio rerio&#039;&#039;) - source: [https://en.wikipedia.org/wiki/Zebrafish Wikipedia] ]]                               &lt;br /&gt;
In this last part of the exercise, we will investigate relationships between different species of fish. We have compiled this list of various fish:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;overflow:auto;&amp;quot;&amp;gt;&lt;br /&gt;
 Latin name             Common name                     TaxID&lt;br /&gt;
 Danio rerio            Zebrafish                        7955 &lt;br /&gt;
 Gadus morhua           Atlantic cod                     8049 &lt;br /&gt;
 Mustelus griseus       Spotless smooth-hound (shark)   89020  &lt;br /&gt;
 Petromyzon marinus     Sea lamprey                      7757 &lt;br /&gt;
 Latimeria chalumnae    Coelacanth (famous &amp;quot;Blue fish&amp;quot;)  7897 &lt;br /&gt;
 Lepidosiren paradoxa   South American lungfish          7883&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
Comparing many species will be tedious by just pointing and clicking, so here we&#039;ll utilize the NCBI &#039;&#039;&#039;Common Tree&#039;&#039;&#039; tool. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Now, go to [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/ the front page of NCBI Taxonomy] and click the &#039;&#039;&#039;Common Tree&#039;&#039;&#039; link a bit down the page (under &amp;quot;Taxonomy resources&amp;quot;). Here, you can add species to a tree one by one by entering either the Latin name or the TaxID in the &amp;quot;&#039;&#039;&#039;search and select&#039;&#039;&#039;&amp;quot; field. (&#039;&#039;It will try to autocomplete the name you&#039;re entering, so pay attention to what is selected when you press enter&#039;&#039;). You can also add a whole list at once if you have a text file containing &#039;&#039;either&#039;&#039; TaxIDs &#039;&#039;or&#039;&#039; Latin names, one per line (doable with a bit of editing of the list above).&lt;br /&gt;
&lt;br /&gt;
[[Image:NcbiCommonTree2026.jpg |400px|frame|center|&#039;&#039;&#039;Direct link:&#039;&#039;&#039; [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/common-tree/ Common Tree tool (2026)] ]]    &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; The 2026 version of the tool is quite good at automatically detecting the shared groups to display, but you can ask the tool to display a full lineage, if needed: You can test this by clicking &amp;quot;&#039;&#039;&#039;View in Taxonomy Browser&#039;&#039;&#039;&amp;quot; after the Common Tree has been populated (but &#039;&#039;&#039;return to the standard view&#039;&#039;&#039; for answering the questions below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4a:&#039;&#039;&#039; In this selection of species, what is the sister group (nearest neighbour) to the Zebrafish? What is the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
Now try to add yourself (i.e. Human) to the tree, using either the Latin name or the TaxID. Any surprises?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4b:&#039;&#039;&#039; What is now the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4c:&#039;&#039;&#039; Which of the following is most closely related to the &amp;quot;Blue fish&amp;quot;: the cod, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4d:&#039;&#039;&#039; Which of the following is most closely related to the cod: the shark, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4e:&#039;&#039;&#039; Which of the following is most closely related to the shark: the lamprey, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4f:&#039;&#039;&#039; Does the category &amp;quot;fish&amp;quot; make any scientific sense?&lt;br /&gt;
&lt;br /&gt;
Leave the browser window with the Common tree open for the next question.&lt;br /&gt;
&lt;br /&gt;
===Comparing trees===&lt;br /&gt;
A bioinformatician has compared the sequences of a gene from the seven species we used in the previous question, and arrived at the following tree:&lt;br /&gt;
&lt;br /&gt;
[[Image:FishNCBI-edited-phy.png]]&lt;br /&gt;
&lt;br /&gt;
You will later learn how to make trees like these in the [[Exercise: Phylogeny|Phylogenetic trees exercise]]. For now, you only need to know that such a tree is not necessarily 100% correct, since it is based on a limited amount of data. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5a:&#039;&#039;&#039; Are there any differences in the branching pattern between the gene tree and the Common tree from the previous question?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5b:&#039;&#039;&#039; Can the gene tree be made to comply with the Common tree by swapping two species? If so, which two?&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
===AI investigation===&lt;br /&gt;
As the final step, use your favorite AI (ChatGPT etc) to try to get as &#039;&#039;&#039;detailed and comprehensive&#039;&#039;&#039; an answer to QUESTION 4f as possible: &#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We do realize that the AI field is rapidly evolving, and answers will vary. We will recommend you to play around with the prompt to make the system generate an answer with the best possible scientific rigor. You are welcome to either use the list of species (+human) as we did above, or ask the AI to find relevant data itself.&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;QUESTION 6:&#039;&#039;&#039; Provide your final prompt + answer from the AI (+ info about which version you used). Reflect upon the correctness of the answer (how much do you trust it, &#039;&#039;anything fishy?&#039;&#039;)&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=968</id>
		<title>Using the Taxonomy database</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=968"/>
		<updated>2026-08-30T19:12:53Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* The NCBI Taxonomy Database */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: &#039;&#039;&#039;Rasmus Wernersson&#039;&#039;&#039; and &#039;&#039;&#039;Henrik Nielsen&#039;&#039;&#039;.&lt;br /&gt;
 &lt;br /&gt;
==Background==&lt;br /&gt;
When comparing DNA and protein sequences from different species it is important to keep in mind that all living organisms at some point in time has shared a common ancestor. Some organisms are closely related and have recently derived from a common ancestor (e.g. Human and Chimpanzee, which diverged 5-10 million years ago) and some are more distantly related (e.g. Human and mouse, diverged 100-150 million years ago).&lt;br /&gt;
&lt;br /&gt;
The more closely related two organisms are, the more similar their sequences will be (say, when comparing the Alpha Globin gene from each of the organisms), and the more likely it will be that similar looking genes from each organism still have the same function (MUCH more about this when we get to pairwise alignment and BLAST searches).&lt;br /&gt;
&lt;br /&gt;
[[file:ApesTax_600.png‎|center|frame|Apes taxonomy (Detailed taxonomy of the Great Apes: Human, Chimp, Gorilla, Orangutan - from Wikipedia)]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Phylogeny vs. taxonomy===&lt;br /&gt;
As we discussed in the lecture all life is organised in a hierarchical taxonomical system, which approximates the &amp;quot;true&amp;quot; underlying phylogeny to a large degree. It&#039;s therefore often important to know where a specific organism is placed in the taxonomical system - this type of information will also always be included along with DNA/Protein sequences from the big databases such as GenBank and UniProt.&lt;br /&gt;
&lt;br /&gt;
Today we will explore various ways to look up and compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
===A word about Wikipedia===&lt;br /&gt;
The free online encyclopedia [http://en.wikipedia.org/wiki/Main_Page Wikipedia] (and other similar resources) is a GREAT way to start out when you need to look up information about a new topic - in this case taxonomy. Almost all species entries in Wikipedia has a &amp;quot;Scientific classification&amp;quot; box which includes taxonomical information (for example see the entry on [http://en.wikipedia.org/wiki/Orangutan Orangutan] or [http://en.wikipedia.org/wiki/Fucus_vesiculosus Fucus vesiculosus] (Bladder wrack / Blæretang)).&lt;br /&gt;
&lt;br /&gt;
HOWEVER: Keep in mind that Wikipedia is NOT a reliable source of information, even if most entries are of a very good quality. The facts in the Wikipedia entries have not been verified by taxonomy experts and can potentially be wrong (everybody can go in and edit the text). We need to look up the taxonomy in an official database (in this case we&#039;ll be using NCBI Taxonomy) before you can state it as a fact.&lt;br /&gt;
&lt;br /&gt;
You CANNOT quote Wikipedia as the only source of your information - you&#039;ll need to find the original primary source of the information or look it up in an official database.&lt;br /&gt;
&lt;br /&gt;
===A word about AI===&lt;br /&gt;
As with Wikipedia using AI to generate a taxonomical analysis (e.g. comparing how a bunch of species are related) can be a great way to get an overview, and a (typically) well written explanation. You will need to ask the AI to include references &#039;&#039;&#039;to actual scientific sources&#039;&#039;&#039; to document the validity of the results, and &#039;&#039;&#039;you will be responsible&#039;&#039;&#039; for double checking the AI output and making sure not only the data is correct, but also that the conclusions are sound.&lt;br /&gt;
&lt;br /&gt;
===The need for a Ground Truth===&lt;br /&gt;
Luckily, modern taxonomy is very well established and there an internationally recognized organization that makes revisions in a highly regulated manner. This &#039;&#039;&#039;taxonomy standard&#039;&#039;&#039; is captured in the &#039;&#039;&#039;NCBI Taxonomy&#039;&#039;&#039; database we&#039;ll be working with in the sections below. For all molecular data NCBI Taxonomy is THE standard reference.&lt;br /&gt;
&lt;br /&gt;
==The NCBI Taxonomy Database==&lt;br /&gt;
[[File:NcbiTax2026.jpg|800px|center|border]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div style=&amp;quot;background-color: lavender; border: solid thin grey;&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;Main link:&#039;&#039;&#039; https://www.ncbi.nlm.nih.gov/datasets/taxonomy/&lt;br /&gt;
&lt;br /&gt;
(&#039;&#039;2026 note: NCBI is in process of migrating to this new site - if you find the old Ncbi Tax homepage via Google, make sure to press the link they provide to jump to the new site&#039;&#039;).&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As mentioned above NCBI Taxonomy will serve as our &#039;&#039;&#039;Ground Truth&#039;&#039;&#039; for everything related to taxonomy. The NCBI Tax database provides the numerical enumeration of species (and other taxonomical levels) that is used and referenced in most Sequence databases, such as GenBank (DNA) and UniProt (Protein). For example human (&#039;&#039;Homo sapiens&#039;&#039;) has the ID &amp;quot;&#039;&#039;&#039;9606&#039;&#039;&#039;&amp;quot; and Yeast (&#039;&#039;Saccharomyces cerevisiae&#039;&#039;) as the ID &amp;quot;&#039;&#039;&#039;4932&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
NCBI Tax is perhaps not a database you would browse for fun (depending on your level of geekiness). It&#039;s good for looking up definitions, and for comparing the taxonomical position of multiple organisms (since the information is so densely presented).&lt;br /&gt;
&lt;br /&gt;
===Example: Homo sapiens===&lt;br /&gt;
# Open the NCBI Taxonomy webpage in a new browser window/tab (see link above)&lt;br /&gt;
# Search for &amp;quot;&#039;&#039;Homo sapiens&#039;&#039;&amp;quot;.&lt;br /&gt;
## Dont Panic: An enormous amount of information is shown - for example about genome sequences. In this case we only need to look at the information presented in the &amp;quot;&#039;&#039;&#039;Taxonomy&#039;&#039;&#039;&amp;quot; tab.&lt;br /&gt;
## Notice the Taxonomy ID - &#039;&#039;&#039;9606&#039;&#039;&#039; as mentioned above.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Lineage&#039;&#039;&#039;&amp;quot; (box on the right hand side): Here a condensed overview of the human lineage is shown by default. Notice the taxonomical ranks we talked about in the lecture (&amp;quot;Phylum&amp;quot;, &amp;quot;Class&amp;quot; etc). You can navigate to the definition of these groups by clicking on them.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Full lineage&#039;&#039;&#039;&amp;quot;: Click this to view a FULL list of all the groups leading &amp;quot;down&amp;quot; to human. Notice, that you can &amp;quot;mouse over&amp;quot; the groups to see a pop-up with the taxonomical rank. NOTICE: The rank &amp;quot;CLADE&amp;quot; is used whenever a group does not have a common English name (&amp;quot;Clade&amp;quot; is the generic name for a uniquely defined taxonomical group).&lt;br /&gt;
&lt;br /&gt;
Play around with the Homo sapiens page for a bit to familiarize yourself with the interface, and answer the following questions along the way:&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1a&#039;&#039;&#039;: What is the TaxID of &amp;quot;&#039;&#039;Metazoa&#039;&#039;&amp;quot;? &lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1b&#039;&#039;&#039;: What is the family that contains humans?&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1c&#039;&#039;&#039;: Are humans vertebrates? (Latin: Vertebrata)?&lt;br /&gt;
&lt;br /&gt;
===Comparing taxonomy using NCBI Tax===&lt;br /&gt;
[[Image:Fruitfly.jpg|thumb|400px|right|Fruit fly (&#039;&#039;Drosophila melanogaster&#039;&#039;) - source: [http://en.wikipedia.org/wiki/Drosophila_melanogaster Wikipedia] ]]Besides being useful for being the official database behind the TaxID&#039;s used in GenBank (and other databases), NCBI Tax actually makes it easy to compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
Let&#039;s take the situation where you have read an interesting paper comparing a DNA sequence between the following three organisms: &#039;&#039;Homo sapiens&#039;&#039; (Human), &#039;&#039;Mus musculus&#039;&#039; (Mouse), and &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly), but you have no idea about the relationship between the three organisms. &lt;br /&gt;
&lt;br /&gt;
We can look this up in NCBI Tax:&lt;br /&gt;
&lt;br /&gt;
# Open two browser windows/tabs and search for &#039;&#039;&#039;Homo sapiens&#039;&#039;&#039; and &#039;&#039;&#039;Mus musculus&#039;&#039;&#039;.&lt;br /&gt;
# By comparing the &amp;quot;lineage&amp;quot; text it will be easy to find out at which taxonomical level human and mouse differ.&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2A&#039;&#039;&#039;: Look at the (simple/abbreviated) lineage information and find lowest ranking common group for human and mouse - what is the name and what is the rank?&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2B&#039;&#039;&#039;: Look at the &amp;quot;&#039;&#039;&#039;full lineage&#039;&#039;&#039; and find the lowest ranking group human and mouse have in common (it&#039;s OK if the rank is &amp;quot;clade&amp;quot;). What is the name and TaxId of the group?&lt;br /&gt;
&lt;br /&gt;
Now repeat the analysis by comparing Human and Fruit Fly (&#039;&#039;Drosophila melanogaster&#039;&#039;)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3A&#039;&#039;&#039;: What is the lowest rank group shared in the (simplified/abbreviated) lineage? (TaxID, Name)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3B&#039;&#039;&#039;: What is the lowest rank group shared in the full lineage? (TaxID, Name)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Fishing in NCBI Tax using the Common Tree function===&lt;br /&gt;
[[Image:Zebrafisch.jpg|thumb|400px|right|Zebrafish (&#039;&#039;Danio rerio&#039;&#039;) - source: [https://en.wikipedia.org/wiki/Zebrafish Wikipedia] ]]                               &lt;br /&gt;
In this last part of the exercise, we will investigate relationships between different species of fish. We have compiled this list of various fish:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;overflow:auto;&amp;quot;&amp;gt;&lt;br /&gt;
 Latin name             Common name                     TaxID&lt;br /&gt;
 Danio rerio            Zebrafish                        7955 &lt;br /&gt;
 Gadus morhua           Atlantic cod                     8049 &lt;br /&gt;
 Mustelus griseus       Spotless smooth-hound (shark)   89020  &lt;br /&gt;
 Petromyzon marinus     Sea lamprey                      7757 &lt;br /&gt;
 Latimeria chalumnae    Coelacanth (famous &amp;quot;Blue fish&amp;quot;)  7897 &lt;br /&gt;
 Lepidosiren paradoxa   South American lungfish          7883&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
Comparing many species will be tedious by just pointing and clicking, so here we&#039;ll utilize the NCBI &#039;&#039;&#039;Common Tree&#039;&#039;&#039; tool. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Now, go to [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/ the front page of NCBI Taxonomy] and click the &#039;&#039;&#039;Common Tree&#039;&#039;&#039; link a bit down the page (under &amp;quot;Taxonomy resources&amp;quot;). Here, you can add species to a tree one by one by entering either the Latin name or the TaxID in the &amp;quot;&#039;&#039;&#039;search and select&#039;&#039;&#039;&amp;quot; field. (&#039;&#039;It will try to autocomplete the name you&#039;re entering, so pay attention to what is selected when you press enter&#039;&#039;). You can also add a whole list at once if you have a text file containing &#039;&#039;either&#039;&#039; TaxIDs &#039;&#039;or&#039;&#039; Latin names, one per line (doable with a bit of editing of the list above).&lt;br /&gt;
&lt;br /&gt;
[[Image:NcbiCommonTree2026.jpg |400px|frame|center|&#039;&#039;&#039;Direct link:&#039;&#039;&#039; [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/common-tree/ Common Tree tool (2026)] ]]    &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; The 2026 version of the tool is quite good at automatically detecting the shared groups to display, but you can ask the tool to display a full lineage, if needed: You can test this by clicking &amp;quot;&#039;&#039;&#039;View in Taxonomy Browser&#039;&#039;&#039;&amp;quot; after the Common Tree has been populated (but &#039;&#039;&#039;return to the standard view&#039;&#039;&#039; for answering the questions below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4a:&#039;&#039;&#039; In this selection of species, what is the sister group (nearest neighbour) to the Zebrafish? What is the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
Now try to add yourself (i.e. Human) to the tree, using either the Latin name or the TaxID. Any surprises?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4b:&#039;&#039;&#039; What is now the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4c:&#039;&#039;&#039; Which of the following is most closely related to the &amp;quot;Blue fish&amp;quot;: the cod, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4d:&#039;&#039;&#039; Which of the following is most closely related to the cod: the shark, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4e:&#039;&#039;&#039; Which of the following is most closely related to the shark: the lamprey, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4f:&#039;&#039;&#039; Does the category &amp;quot;fish&amp;quot; make any scientific sense?&lt;br /&gt;
&lt;br /&gt;
Leave the browser window with the Common tree open for the next question.&lt;br /&gt;
&lt;br /&gt;
===Comparing trees===&lt;br /&gt;
A bioinformatician has compared the sequences of a gene from the seven species we used in the previous question, and arrived at the following tree:&lt;br /&gt;
&lt;br /&gt;
[[Image:FishNCBI-edited-phy.png]]&lt;br /&gt;
&lt;br /&gt;
You will later learn how to make trees like these in the [[Exercise: Phylogeny|Phylogenetic trees exercise]]. For now, you only need to know that such a tree is not necessarily 100% correct, since it is based on a limited amount of data. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5a:&#039;&#039;&#039; Are there any differences in the branching pattern between the gene tree and the Common tree from the previous question?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5b:&#039;&#039;&#039; Can the gene tree be made to comply with the Common tree by swapping two species? If so, which two?&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
===AI investigation===&lt;br /&gt;
As the final step, use your favorite AI (ChatGPT etc) to try to get as &#039;&#039;&#039;detailed and comprehensive&#039;&#039;&#039; an answer to QUESTION 4f as possible: &#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We do realize that the AI field is rapidly evolving, and answers will vary. We will recommend you to play around with the prompt to make the system generate an answer with the best possible scientific rigor. You are welcome to either use the list of species (+human) as we did above, or ask the AI to find relevant data itself.&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;QUESTION 6:&#039;&#039;&#039; Provide your final prompt + answer from the AI (+ info about which version you used). Reflect upon the correctness of the answer (how much do you trust it, &#039;&#039;anything fishy?&#039;&#039;)&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=920</id>
		<title>Using the Taxonomy database</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database&amp;diff=920"/>
		<updated>2026-08-28T12:30:14Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Exercise written by: &#039;&#039;&#039;Rasmus Wernersson&#039;&#039;&#039; and &#039;&#039;&#039;Henrik Nielsen&#039;&#039;&#039;.&lt;br /&gt;
 &lt;br /&gt;
==Background==&lt;br /&gt;
When comparing DNA and protein sequences from different species it is important to keep in mind that all living organisms at some point in time has shared a common ancestor. Some organisms are closely related and have recently derived from a common ancestor (e.g. Human and Chimpanzee, which diverged 5-10 million years ago) and some are more distantly related (e.g. Human and mouse, diverged 100-150 million years ago).&lt;br /&gt;
&lt;br /&gt;
The more closely related two organisms are, the more similar their sequences will be (say, when comparing the Alpha Globin gene from each of the organisms), and the more likely it will be that similar looking genes from each organism still have the same function (MUCH more about this when we get to pairwise alignment and BLAST searches).&lt;br /&gt;
&lt;br /&gt;
[[file:ApesTax_600.png‎|center|frame|Apes taxonomy (Detailed taxonomy of the Great Apes: Human, Chimp, Gorilla, Orangutan - from Wikipedia)]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Phylogeny vs. taxonomy===&lt;br /&gt;
As we discussed in the lecture all life is organised in a hierarchical taxonomical system, which approximates the &amp;quot;true&amp;quot; underlying phylogeny to a large degree. It&#039;s therefore often important to know where a specific organism is placed in the taxonomical system - this type of information will also always be included along with DNA/Protein sequences from the big databases such as GenBank and UniProt.&lt;br /&gt;
&lt;br /&gt;
Today we will explore various ways to look up and compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
===A word about Wikipedia===&lt;br /&gt;
The free online encyclopedia [http://en.wikipedia.org/wiki/Main_Page Wikipedia] (and other similar resources) is a GREAT way to start out when you need to look up information about a new topic - in this case taxonomy. Almost all species entries in Wikipedia has a &amp;quot;Scientific classification&amp;quot; box which includes taxonomical information (for example see the entry on [http://en.wikipedia.org/wiki/Orangutan Orangutan] or [http://en.wikipedia.org/wiki/Fucus_vesiculosus Fucus vesiculosus] (Bladder wrack / Blæretang)).&lt;br /&gt;
&lt;br /&gt;
HOWEVER: Keep in mind that Wikipedia is NOT a reliable source of information, even if most entries are of a very good quality. The facts in the Wikipedia entries have not been verified by taxonomy experts and can potentially be wrong (everybody can go in and edit the text). We need to look up the taxonomy in an official database (in this case we&#039;ll be using NCBI Taxonomy) before you can state it as a fact.&lt;br /&gt;
&lt;br /&gt;
You CANNOT quote Wikipedia as the only source of your information - you&#039;ll need to find the original primary source of the information or look it up in an official database.&lt;br /&gt;
&lt;br /&gt;
===A word about AI===&lt;br /&gt;
As with Wikipedia using AI to generate a taxonomical analysis (e.g. comparing how a bunch of species are related) can be a great way to get an overview, and a (typically) well written explanation. You will need to ask the AI to include references &#039;&#039;&#039;to actual scientific sources&#039;&#039;&#039; to document the validity of the results, and &#039;&#039;&#039;you will be responsible&#039;&#039;&#039; for double checking the AI output and making sure not only the data is correct, but also that the conclusions are sound.&lt;br /&gt;
&lt;br /&gt;
===The need for a Ground Truth===&lt;br /&gt;
Luckily, modern taxonomy is very well established and there an internationally recognized organization that makes revisions in a highly regulated manner. This &#039;&#039;&#039;taxonomy standard&#039;&#039;&#039; is captured in the &#039;&#039;&#039;NCBI Taxonomy&#039;&#039;&#039; database we&#039;ll be working with in the sections below. For all molecular data NCBI Taxonomy is THE standard reference.&lt;br /&gt;
&lt;br /&gt;
==The NCBI Taxonomy Database==&lt;br /&gt;
[[File:NcbiTax2026.jpg|800px|center|border]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Main link:&#039;&#039;&#039; https://www.ncbi.nlm.nih.gov/datasets/taxonomy/&lt;br /&gt;
&lt;br /&gt;
(&#039;&#039;2026 note: NCBI is in process of migrating to this new site - if you find the old Ncbi Tax homepage via Google, make sure to press the link they provide to jump to the new site&#039;&#039;).&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
As mentioned above NCBI Taxonomy will serve as our &#039;&#039;&#039;Ground Truth&#039;&#039;&#039; for everything related to taxonomy. The NCBI Tax database provides the numerical enumeration of species (and other taxonomical levels) that is used and referenced in most Sequence databases, such as GenBank (DNA) and UniProt (Protein). For example human (&#039;&#039;Homo sapiens&#039;&#039;) has the ID &amp;quot;&#039;&#039;&#039;9606&#039;&#039;&#039;&amp;quot; and Yeast (&#039;&#039;Saccharomyces cerevisiae&#039;&#039;) as the ID &amp;quot;&#039;&#039;&#039;4932&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
NCBI Tax is perhaps not a database you would browse for fun (depending on your level of geekiness). It&#039;s good for looking up definitions, and for comparing the taxonomical position of multiple organisms (since the information is so densely presented).&lt;br /&gt;
&lt;br /&gt;
===Example: Homo sapiens===&lt;br /&gt;
# Open the NCBI Taxonomy webpage in a new browser window/tab (see link above)&lt;br /&gt;
# Search for &amp;quot;&#039;&#039;Homo sapiens&#039;&#039;&amp;quot;.&lt;br /&gt;
## Dont Panic: An enormous amount of information is shown - for example about genome sequences. In this case we only need to look at the information presented in the &amp;quot;&#039;&#039;&#039;Taxonomy&#039;&#039;&#039;&amp;quot; tab.&lt;br /&gt;
## Notice the Taxonomy ID - &#039;&#039;&#039;9606&#039;&#039;&#039; as mentioned above.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Lineage&#039;&#039;&#039;&amp;quot; (box on the right hand side): Here a condensed overview of the human lineage is shown by default. Notice the taxonomical ranks we talked about in the lecture (&amp;quot;Phylum&amp;quot;, &amp;quot;Class&amp;quot; etc). You can navigate to the definition of these groups by clicking on them.&lt;br /&gt;
## &amp;quot;&#039;&#039;&#039;Full lineage&#039;&#039;&#039;&amp;quot;: Click this to view a FULL list of all the groups leading &amp;quot;down&amp;quot; to human. Notice, that you can &amp;quot;mouse over&amp;quot; the groups to see a pop-up with the taxonomical rank. NOTICE: The rank &amp;quot;CLADE&amp;quot; is used whenever a group does not have a common English name (&amp;quot;Clade&amp;quot; is the generic name for a uniquely defined taxonomical group).&lt;br /&gt;
&lt;br /&gt;
Play around with the Homo sapiens page for a bit to familiarize you with the interface, and answer the following questions along the way:&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1a&#039;&#039;&#039;: What is the TaxID of &amp;quot;&#039;&#039;Metazoa&#039;&#039;&amp;quot;? &lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1b&#039;&#039;&#039;: What is the family that contains humans?&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 1c&#039;&#039;&#039;: Are humans vertebrates? (Latin: Vertebrata)?&lt;br /&gt;
&lt;br /&gt;
===Comparing taxonomy using NCBI Tax===&lt;br /&gt;
[[Image:Fruitfly.jpg|thumb|400px|right|Fruit fly (&#039;&#039;Drosophila melanogaster&#039;&#039;) - source: [http://en.wikipedia.org/wiki/Drosophila_melanogaster Wikipedia] ]]Besides being useful for being the official database behind the TaxID&#039;s used in GenBank (and other databases), NCBI Tax actually makes it easy to compare taxonomy.&lt;br /&gt;
&lt;br /&gt;
Let&#039;s take the situation where you have read an interesting paper comparing a DNA sequence between the following three organisms: &#039;&#039;Homo sapiens&#039;&#039; (Human), &#039;&#039;Mus musculus&#039;&#039; (Mouse), and &#039;&#039;Drosophila melanogaster&#039;&#039; (Fruit fly), but you have no idea about the relationship between the three organisms. &lt;br /&gt;
&lt;br /&gt;
We can look this up in NCBI Tax:&lt;br /&gt;
&lt;br /&gt;
# Open two browser windows/tabs and search for &#039;&#039;&#039;Homo sapiens&#039;&#039;&#039; and &#039;&#039;&#039;Mus musculus&#039;&#039;&#039;.&lt;br /&gt;
# By comparing the &amp;quot;lineage&amp;quot; text it will be easy to find out at which taxonomical level human and mouse differ.&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2A&#039;&#039;&#039;: Look at the (simple/abbreviated) lineage information and find lowest ranking common group for human and mouse - what is the name and what is the rank?&lt;br /&gt;
# &#039;&#039;&#039;QUESTION 2B&#039;&#039;&#039;: Look at the &amp;quot;&#039;&#039;&#039;full lineage&#039;&#039;&#039; and find the lowest ranking group human and mouse have in common (it&#039;s OK if the rank is &amp;quot;clade&amp;quot;). What is the name and TaxId of the group?&lt;br /&gt;
&lt;br /&gt;
Now repeat the analysis by comparing Human and Fruit Fly (&#039;&#039;Drosophila melanogaster&#039;&#039;)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3A&#039;&#039;&#039;: What is the lowest rank group shared in the (simplified/abbreviated) lineage? (TaxID, Name)&lt;br /&gt;
* &#039;&#039;&#039;QUESTION 3B&#039;&#039;&#039;: What is the lowest rank group shared in the full lineage? (TaxID, Name)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Fishing in NCBI Tax using the Common Tree function===&lt;br /&gt;
[[Image:Zebrafisch.jpg|thumb|400px|right|Zebrafish (&#039;&#039;Danio rerio&#039;&#039;) - source: [https://en.wikipedia.org/wiki/Zebrafish Wikipedia] ]]                               &lt;br /&gt;
In this last part of the exercise, we will investigate relationships between different species of fish. We have compiled this list of various fish:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;overflow:auto;&amp;quot;&amp;gt;&lt;br /&gt;
 Latin name             Common name                     TaxID&lt;br /&gt;
 Danio rerio            Zebrafish                        7955 &lt;br /&gt;
 Gadus morhua           Atlantic cod                     8049 &lt;br /&gt;
 Mustelus griseus       Spotless smooth-hound (shark)   89020  &lt;br /&gt;
 Petromyzon marinus     Sea lamprey                      7757 &lt;br /&gt;
 Latimeria chalumnae    Coelacanth (famous &amp;quot;Blue fish&amp;quot;)  7897 &lt;br /&gt;
 Lepidosiren paradoxa   South American lungfish          7883&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
Comparing many species will be tedious by just pointing and clicking, so here we&#039;ll utilize the NCBI &#039;&#039;&#039;Common Tree&#039;&#039;&#039; tool. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Now, go to [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/ the front page of NCBI Taxonomy] and click the &#039;&#039;&#039;Common Tree&#039;&#039;&#039; link a bit down the page (under &amp;quot;Taxonomy resources&amp;quot;). Here, you can add species to a tree one by one by entering either the Latin name or the TaxID in the &amp;quot;&#039;&#039;&#039;search and select&#039;&#039;&#039;&amp;quot; field. (&#039;&#039;It will try to autocomplete the name you&#039;re entering, so pay attention to what is selected when you press enter&#039;&#039;). You can also add a whole list at once if you have a text file containing &#039;&#039;either&#039;&#039; TaxIDs &#039;&#039;or&#039;&#039; Latin names, one per line (doable with a bit of editing of the list above).&lt;br /&gt;
&lt;br /&gt;
[[Image:NcbiCommonTree2026.jpg |400px|frame|center|&#039;&#039;&#039;Direct link:&#039;&#039;&#039; [https://www.ncbi.nlm.nih.gov/datasets/taxonomy/common-tree/ Common Tree tool (2026)] ]]    &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;IMPORTANT:&#039;&#039;&#039; The 2026 version of the tool is quite good at automatically detecting the shared groups to display, but you can ask the tool to display a full lineage, if needed: You can test this by clicking &amp;quot;&#039;&#039;&#039;View in Taxonomy Browser&#039;&#039;&#039;&amp;quot; after the Common Tree has been populated (but &#039;&#039;&#039;return to the standard view&#039;&#039;&#039; for answering the questions below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4a:&#039;&#039;&#039; In this selection of species, what is the sister group (nearest neighbour) to the Zebrafish? What is the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
Now try to add yourself (i.e. Human) to the tree, using either the Latin name or the TaxID. Any surprises?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4b:&#039;&#039;&#039; What is now the sister group to the lungfish?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4c:&#039;&#039;&#039; Which of the following is most closely related to the &amp;quot;Blue fish&amp;quot;: the cod, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4d:&#039;&#039;&#039; Which of the following is most closely related to the cod: the shark, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4e:&#039;&#039;&#039; Which of the following is most closely related to the shark: the lamprey, or you?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 4f:&#039;&#039;&#039; Does the category &amp;quot;fish&amp;quot; make any scientific sense?&lt;br /&gt;
&lt;br /&gt;
Leave the browser window with the Common tree open for the next question.&lt;br /&gt;
&lt;br /&gt;
===Comparing trees===&lt;br /&gt;
A bioinformatician has compared the sequences of a gene from the seven species we used in the previous question, and arrived at the following tree:&lt;br /&gt;
&lt;br /&gt;
[[Image:FishNCBI-edited-phy.png]]&lt;br /&gt;
&lt;br /&gt;
You will later learn how to make trees like these in the [[Exercise: Phylogeny|Phylogenetic trees exercise]]. For now, you only need to know that such a tree is not necessarily 100% correct, since it is based on a limited amount of data. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5a:&#039;&#039;&#039; Are there any differences in the branching pattern between the gene tree and the Common tree from the previous question?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;QUESTION 5b:&#039;&#039;&#039; Can the gene tree be made to comply with the Common tree by swapping two species? If so, which two?&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
===AI investigation===&lt;br /&gt;
As the final step, use your favorite AI (ChatGPT etc) to try to get as &#039;&#039;&#039;detailed and comprehensive&#039;&#039;&#039; an answer to QUESTION 4f as possible: &#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We do realize that the AI field is rapidly evolving, and answers will vary. We will recommend you to play around with the prompt to make the system generate an answer with the best possible scientific rigor. You are welcome to either use the list of species (+human) as we did above, or ask the AI to find relevant data itself.&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;QUESTION 6:&#039;&#039;&#039; Provide your final prompt + answer from the AI (+ info about which version you used). Reflect upon the correctness of the answer (how much do you trust it, &#039;&#039;anything fishy?&#039;&#039;)&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=919</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=919"/>
		<updated>2026-08-28T10:18:27Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sits between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px|left]]&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
We can ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=918</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=918"/>
		<updated>2026-08-28T10:18:03Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sits between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px|left]]&lt;br /&gt;
&amp;lt;br style=&amp;quot;clear: both&amp;quot; /&amp;gt;&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=917</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=917"/>
		<updated>2026-08-28T10:17:01Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sits between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px|left]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=916</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=916"/>
		<updated>2026-08-28T10:16:11Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 3: */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sits between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=915</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=915"/>
		<updated>2026-08-28T10:15:48Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 1: */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrate (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=914</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=914"/>
		<updated>2026-08-28T10:12:26Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Answers to the Using the Taxonomy database exercise */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson and Henrik Nielsen. Last update: Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrated (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=913</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=913"/>
		<updated>2026-08-28T10:10:21Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson, Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrated (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Alert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=912</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=912"/>
		<updated>2026-08-28T10:09:23Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson, Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrated (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Atlert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below. It is certainly possible to write it in a more concise manner, but it does capture the main points, and includes references beyond popular science articles and wikipedia. Notice that part 2 specifically talks about the distance in relationship between several of the brances of the tree we also have in our analysis (we use Cod, the analysis here uses Salmon as an example - both cases includes human and the lungfish).&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=911</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=911"/>
		<updated>2026-08-28T10:04:09Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson, Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrated (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Atlert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the parts below.&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
[[Image: Taxonomy_AI_answer_2026_extC.jpg|frame|left]]&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Taxonomy_AI_answer_2026_extC.jpg&amp;diff=910</id>
		<title>File:Taxonomy AI answer 2026 extC.jpg</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Taxonomy_AI_answer_2026_extC.jpg&amp;diff=910"/>
		<updated>2026-08-28T10:03:38Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=909</id>
		<title>Using the Taxonomy database ANSWERS</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=Using_the_Taxonomy_database_ANSWERS&amp;diff=909"/>
		<updated>2026-08-28T10:02:39Z</updated>

		<summary type="html">&lt;p&gt;Raz: /* Question 6: (AI assisted analysis) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Answers to the [[Using_the_Taxonomy_database|Using the Taxonomy database exercise]]=&lt;br /&gt;
&lt;br /&gt;
Answers by: Rasmus Wernersson, Aug 2026&lt;br /&gt;
 &lt;br /&gt;
== Question 1:==&lt;br /&gt;
*&#039;&#039;&#039;1a:&#039;&#039;&#039; &amp;quot;Metazoa&amp;quot; Taxonomy ID: 33208&lt;br /&gt;
*&#039;&#039;&#039;1b:&#039;&#039;&#039; Familiy containing humans: &#039;&#039;Hominidae&#039;&#039;&lt;br /&gt;
*&#039;&#039;&#039;1c:&#039;&#039;&#039; Yes, humans are certainly vertebrated (we have a spine), and that can been seen in the taxonomy by looking at the &amp;quot;full lineage&amp;quot; and seing we&#039;re members of the group &amp;quot;Vertebrata&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Question 2: ==&lt;br /&gt;
&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Mouse (TaxId:10090)&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;2a:&#039;&#039;&#039; (abbreviated lineage): &#039;&#039;&#039;Mammalia&#039;&#039;&#039; (rank: class) - Eng: Mammals&lt;br /&gt;
*&#039;&#039;&#039;2b:&#039;&#039;&#039; (full lineage): &#039;&#039;&#039;Euarchontoglires&#039;&#039;&#039; (rank: superorder) is the last common group before human and mouse branch off into &#039;&#039;&#039;primates&#039;&#039;&#039; and &#039;&#039;&#039;rodents&#039;&#039;&#039;. This is a group of placental mammals.&lt;br /&gt;
&lt;br /&gt;
== Question 3: ==&lt;br /&gt;
Lowest ranking taxonomical group shared by Human (TaxID: 9606) and Fruit Fly (&#039;&#039;D. melanogaster&#039;&#039; - TaxId:7227) &lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;3a:&#039;&#039;&#039; (Abbreviated lineage): &#039;&#039;&#039;Metazoa&#039;&#039;&#039; (rank: kingdom) - this is the group of all animals.&lt;br /&gt;
*&#039;&#039;&#039;3b:&#039;&#039;&#039; (Full lineage): &#039;&#039;&#039;Bilateria&#039;&#039;&#039; (rank: clade (which just mean a group) - sit between kingdom (Metazoa) and phyllum (&#039;&#039;Cordata&#039;&#039; for human, &#039;&#039;Arthropoda&#039;&#039; for the fly))&lt;br /&gt;
&lt;br /&gt;
== Question 4: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
The sister group to the Zebrafish is the cod.&lt;br /&gt;
&lt;br /&gt;
The sister group to the lungfish is the Coelacanth (famous &amp;quot;Blue fish&amp;quot;).&lt;br /&gt;
 &lt;br /&gt;
=== b) ===&lt;br /&gt;
Now, the sister group to the lungfish is you!&lt;br /&gt;
 &lt;br /&gt;
=== c) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== d) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== e) ===&lt;br /&gt;
you&lt;br /&gt;
 &lt;br /&gt;
=== f) ===&lt;br /&gt;
No! See &lt;br /&gt;
https://www.sciencealert.com/actually-there-is-no-such-thing-as-a-fish-say-cladists&lt;br /&gt;
&lt;br /&gt;
== Question 5: ==&lt;br /&gt;
&lt;br /&gt;
===a) ===&lt;br /&gt;
Yes, there is a difference.&lt;br /&gt;
=== b) ===&lt;br /&gt;
Yes, the trees can be made identical by swapping the Coelacanth and the lungfish.&lt;br /&gt;
&lt;br /&gt;
== Question 6: (AI assisted analysis) ==&lt;br /&gt;
&lt;br /&gt;
No single answer can be given here - we&#039;ll discuss some example in class next week as part of the wrap-up.&lt;br /&gt;
&lt;br /&gt;
In august 2026 a simple query Google AI with the question &amp;quot;&#039;&#039;Does the category &amp;quot;fish&amp;quot; make any scientific sense? &#039;&#039;&amp;quot; gives a rather good overview explanation (see screenshot below).&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTICE:&#039;&#039;&#039; It&#039;s clear that the AI has sourced (some of) the information from the Science Atlert article we have also given as a reference under the answer to &#039;&#039;&#039;Q4f&#039;&#039;&#039; above.&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026.jpg|thumb|1000px]]&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
If we ask the AI to dig deeper and add more scientific rigor by adding the prompt:&lt;br /&gt;
&lt;br /&gt;
 Follow up by expanding the answer with more scientific rigor. &lt;br /&gt;
 Cite real scientific sources. This can both be scientific publications as well as &lt;br /&gt;
 scientific reference databases (for example, but not limited to, NCBI Taxonomy).&lt;br /&gt;
&lt;br /&gt;
This gives an answer broken down into the following parts:&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extA.jpg|frame|left]]&lt;br /&gt;
xxxx&lt;br /&gt;
&lt;br /&gt;
[[Image:Taxonomy_AI_answer_2026_extB.jpg|frame|left]]&lt;br /&gt;
xxxx&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Taxonomy_AI_answer_2026_extB.jpg&amp;diff=908</id>
		<title>File:Taxonomy AI answer 2026 extB.jpg</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk/22111/index.php?title=File:Taxonomy_AI_answer_2026_extB.jpg&amp;diff=908"/>
		<updated>2026-08-28T10:02:18Z</updated>

		<summary type="html">&lt;p&gt;Raz: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Raz</name></author>
	</entry>
</feed>