For this exercise we'll use a tool specifically build for experimentation with LOGO plots and for building understanding of how they work internally.
[[image:LogoLab_Sep2026.png|500px|right]]
For this exercise we'll use a tool specifically built for experimentation with LOGO plots and for building understanding of how they work internally.
'''Link:''' https://wernersson.dk/logolab/ (Build by Rasmus Wernersson, it'll migrate to a DTU server soon).
'''Link:''' https://wernersson.dk/logolab/ (Built by Rasmus Wernersson, it'll migrate to a DTU server soon).
'''Points of notice:'''
'''Points of notice:'''
* It runs entirely in your browser, and allows for very quick experimentarion with the visualization.
* It runs entirely in your browser, and allows for very quick experimentation with the visualization.
* The is a stand-alone downloadable build available for running locally on you own machine.
* There is a stand-alone downloadable build available for running locally on your own machine.
* It comes with an "Inspector" functionality that allows you to "peak inside" each position in the LOGO to see the numbers behind the visuals.
* It comes with an "Inspector" functionality that allows you to "peek inside" each position in the LOGO to see the numbers behind the visuals.
* It has an extensive help-page going over all the equations used.
* It has an extensive help-page going over all the equations used.
It you want a good command-line tool as well, we'll recommend the "weblogo" program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/
It you want a good command-line tool as well, we'll recommend the "weblogo" program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/
<br style="clear: both" />
== Seq2Logo (CBS/DTU) ==
== Seq2Logo (CBS/DTU) ==
Line 21:
Line 23:
A more advanced method for working with '''peptide''' sequences.
A more advanced method for working with '''peptide''' sequences.
We'll introduce it briefly here, and get back to it in more details when we get to work with weight matrices.
<br style="clear: both" />
= Part 0: Getting to know "LogoLab" =
= Part 0: Getting to know "LogoLab" =
Line 32:
Line 37:
With the DNA example do the following:
With the DNA example do the following:
# Play around with the '''logo model''' - make sure you understand what each model is about (cross-ref with the lecture slides).
# Play around with the '''logo model''' - make sure you understand what each model is about (cross-ref with the lecture slides).
# Return to "Shannon information" as the model, and observer what happens in the "Inspector" view (scroll down a little bit to see it) when you click a position in the LOGO.
# Return to "Shannon information" as the model, and observe what happens in the "Inspector" view (scroll down a little bit to see it) when you click a position in the LOGO.
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.
# Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.
# Add error-bars to your plot and, play around with the "Display option" to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that '''you can set a title for you plot''' using the display options.
# Add error-bars to your plot and, play around with the "Display option" to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that '''you can set a title for your plot''' using the display options.
== Investigating the peptide example ==
== Investigating the peptide example ==
Line 42:
Line 47:
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.
# As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.
# With the KL model active, click though some of the positions where many AA's are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.
# With the KL model active, click though some of the positions where many AA's are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.
# Pseudo-counts: we'll learn much more about them next week, but they are essentially a way have a best guess of amino-acid frequecies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on anf off to see (visually) what that do - I'll return to them next week.
# Pseudo-counts: we'll learn much more about them later in the course, but they are essentially a way to have a best guess of amino-acid frequencies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on and off to see (visually) what that do - I'll return to them later in the course.
= Part 1: DNA logos (scientific case) =
= Part 1: DNA logos (scientific case) =
Line 213:
Line 218:
=== Signal peptide comparison ===
=== Signal peptide comparison ===
'''OVERALL TASK:'''
'''OVERALL TASK:'''
* Generating peptide LOGOs with the WebLogo resource is as straightforward as before, and your goal is to '''compare''' the motifs in the signal peptides across the three taxonomical groups.
* Generating peptide LOGOs with the Logo Lab resource is as straightforward as before, and your goal is to '''compare''' the motifs in the signal peptides across the three taxonomical groups.
* Play around with the options (dimensions, Y-axis scaling, tics etc) as much as you like, but '''make sure that''':
* Play around with the options (dimensions etc) as much as you like, but '''make sure that''':
*# It's easy to compare the plots (put in labels etc)
*# It's easy to compare the plots (put in titles etc)
*# The y-axis scaling must be the same across the 3 plots
<!-- *# The y-axis scaling must be the same across the 3 plots -->
*# Plot the '''signal peptide part''' of the sequences in '''negative numbers'''
*# Plot the '''signal peptide part''' of the sequences in '''negative numbers'''
*# Read up on what the coloring means (click the "?" in "color scheme) - this will help you interpret the results.
*# Read up on what the "Chemistry" coloring means (see below the plot) - this will help you interpret the results.
'''2026 note:''' The new LogoLab webapp also provides the needed functionality for this part, but it's a little less tested on KL logos. Once you have done the part below with seq2logo (which we'll also be using next week) you're welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.
'''2026 note:''' The new LogoLab webapp also provides the needed functionality for this part, but it's a little less tested on KL logos. Once you have done the part below with seq2logo (which we'll also be using while working with weight matrices) you're welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.
</div>
</div>
As the next step we shall investigate how to work with the '''Seq2Logo''' tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we'll get back to next week when we work with '''weight matrices'''.
As the next step we shall investigate how to work with the '''Seq2Logo''' tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we'll get back to later in the course when we work with '''weight matrices'''.
'''TASK:''' create a basic LOGO plot for the '''eukaryote''' dataset
'''TASK:''' create a basic LOGO plot for the '''eukaryote''' dataset
In this exercise we will introduce two methods for generating sequences logos, and we will investigate, how we can extract and compare sequences information from large sets of sequences.
LogoLab (DTU)
For this exercise we'll use a tool specifically built for experimentation with LOGO plots and for building understanding of how they work internally.
It runs entirely in your browser, and allows for very quick experimentation with the visualization.
There is a stand-alone downloadable build available for running locally on your own machine.
It comes with an "Inspector" functionality that allows you to "peek inside" each position in the LOGO to see the numbers behind the visuals.
It has an extensive help-page going over all the equations used.
It you want a good command-line tool as well, we'll recommend the "weblogo" program (Berkeley, not DTU) - it can be downloaded here: http://weblogo.berkeley.edu/
The app comes with predefined examples of DNA and peptide alignments - those are purely made up (synthetic) examples with the purpose of high-lighting how the tool works.
With the DNA example do the following:
Play around with the logo model - make sure you understand what each model is about (cross-ref with the lecture slides).
Return to "Shannon information" as the model, and observe what happens in the "Inspector" view (scroll down a little bit to see it) when you click a position in the LOGO.
Make sure you understand why the 1st position has an information content of 2.0 bits, and the 3rd position has an information content of 1.0 bits. Make sure you understand the scaling (heights) of the letters - think about it in the context of the hand-out exercise we did together during the lecture.
Add error-bars to your plot and, play around with the "Display option" to do on-the-fly rescaling of the LOGO plot, and test the option for downloading the plot to you own computer (PNG is likely the easiest to use). Notice that you can set a title for your plot using the display options.
Investigating the peptide example
Load the peptide example (click the "protein" button), and play around with the logo much like in the DNA example:
Try out the different coloring schemes, and make sure you understand what amino-acid properties they color code.
As we talked about in class, background correction is much more important for protein sequences. Try out the different logo models, and notice that the KL model can make a substantial difference.
With the KL model active, click though some of the positions where many AA's are seen in the data, and look at the data in the Inspector view. Make sure you understand where p and q comes from, and how they affect the calculations.
Pseudo-counts: we'll learn much more about them later in the course, but they are essentially a way to have a best guess of amino-acid frequencies we might see if we had a larger data set (including the amino-acids we have not seen in out data). Try to briefly click them on and off to see (visually) what that do - I'll return to them later in the course.
Part 1: DNA logos (scientific case)
We'll start out by investigating a small dataset of human splice sites (donor/acceptor)
HUMAN donor sites
The sequences below has been extracted from a random sample of human genes, and each line corresponds to the DNA sequence immetidately BEFORE and AFTER the EXON/INTRON boundary. You can think of it as a multiple alignment written in a compact way, where all the sequence names have been discarded.
Use the LogoLab webapp to create a new logo using the sequences above. Notice that you can clear the example sequences and paste in your own.
QUESTION #1:
Paste in the logo in your report (PNG download is an easy way to get it).
Can you recognize the DONOR site pattern? (Compare to the lecture slides) - how many bits of information are in the GT positions?
How many bases are from the EXON and how many are from the INTRON?
TASK 2: prettifying the LOGO
Since the interesting part of the sequences is at the EXON/INTRON boundary, it would be nice to be able to high-light this. An easy way to achieve this is to adjust the numbering scheme to display the GT to start at position "0".
Play around with the "Numbering starts at" setting to give the EXON sequence negative numbers and the INTRON sequence positive numbers.
QUESTION #2:
Add a title to you plot using the "display options" and include it your report.
TASK/QUESTION #3:
Finally for good measure we also want to generate a frequency plot for the human donor sites.
Find the option to do this, and paste in the resultant LOGO in your report.
Research task: cross-species comparison
As we have seen above, it's pretty straightforward to generate DNA logos (provided that the data has already been well prepared), and it's now time to perform a real research task.
INPUT DATA:
The ZIP file linked below, contains 500 sequences related to DONOR and ACCEPTOR sites for the following species:
It's is now your task to investigate the signal around first the DONOR and then the ACCEPTOR site for all 5 species, and conclude what is similar and what is different.
IMPORTANT:
Make sure to label all your plot with species + donor/acceptor (you will be generating 10 different plots, and it's important that they are easy to tell apart).
QUESTION #4:
Include all LOGOs in your report, and note down your observations — e.g. if less signal is seen in the EXON part, will more signal be seen in the INTRON part?
Looking for much weaker DNA motifs
In the examples above, we have been investigating some pretty strong signals, and now we turn our attention to how to work with data sets with somewhat weaker motifs.
Shine-Dalgarno sequence
In prokaryotes the translation of a mRNA transcript is initiated by the binding of the ribosome to the mRNA a little upstream of the start codon. This binding site is known as the RBS (ribosomal binding site) and the sequence being recognized by the ribosome is known as the "Shine-Dalgarno sequence". The consensus sequence is AGGAGG (DNA: AGGAGG) and it's located approximately 8bp before the AUG (DNA: ATG).
In order to investigate this in more details, we have prepared a data set of 500 E. coli genes which includes 50 bp before the START codon and the first 50 bp of the coding sequence (100 bp in all).
Generate a LOGO of the 500 E. coli sequences. Since we're visualizing 100bps we need to find a good balance with height and width of the plot - play around with the Display options to produce a plot that also looks fine when exported to a PNG.
As it can easily be seen, much of the sequence positions contain next to no signal, and we can benefit from narrowing down the region we look at.
TASK:
Generate a new LOGO of the E. coli data set, but set the Alignment range (found below the sequence input text field) to position 25-75 to only show the middle part of the data.
Play around with the dimensions to get a good looking LOGO.
QUESTION #5:
Paste your LOGO into the report.
Is the start codon always "ATG"?
Can you see something resembling the Shine-Dalgarno sequence (at which positions)?
In order to get a better "view" of the sequence region with the Shine-Dalgarno sequence, we need to zoom in a bit further. The very strong signal from the START codon is blinding us a bit, and it can be beneficial to remove that from our field of view.
TASK:
Set the Alignment range to position 30-50, and play around with the dimensions to get a good looking plot.
QUESTION #6:
Paste your LOGO into the report.
Is the logo (more or less) consistent with the 6 bp consensus sequence?
Kozak sequence
As the final step in our work with DNA logos, we shall investigate the signal around eukaryotic START codons, and see if we can find the "Kozak sequence" which aids in the initiation of eukaryotic translation. For this we have prepared a data set of 500 yeast sequences (50 bp before+after CDS start - exactly as above).
It's is now your task to investigate, if you can find a signal upstream (before) the START codon, and prepare a good visualization of this using the tricks you have learned so far.
Can you find any positions (except for the ATG) with information content above 0.2 bits?
For any such position what is the approximate frequency of the most common base?
Include relevant LOGO plots in your report
Part 2: Protein logos
We'll start our work with peptide LOGOs with a study of signal peptides, which is a well understood system, and we have prepared data files for this in advance.
Most secreted proteins have an N-terminal signal peptide that directs the protein to the secretory pathway. During passage of the membrane (plasma membrane in prokaryotes, ER membrane in eukaryotes), the signal peptide is cleaved off.
Signal peptides are known from all domains of life, but there are certain differences between signal peptides of various taxonomical groups. Eukaryotic signal peptides are on average shorter than those of bacteria, and Gram-negative bacteria (bacteria with two membranes) have shorter signal peptides than Gram-positive bacteria (bacteria with one membrane). (Not much is known about Archaeal signal peptides).
It is now your task to investigate whether there are other differences between signal peptides of Eukaryotes and the two bacterial groups.
DATA: Zip archive signal peptides for all 3 taxonomical groups. Download link:signal_peptides.zip
#Seqs #File name
3280 EUK.sp.25+5.fasta
416 gram+.sp.25+5.fasta
846 gram-.sp.25+5.fasta
Each sequence line contains (up to) 25aa signal peptide + 5aa following the cleavage site. All sequences has been aligned at the cleavage site as illustrated below. Notice that not all signal peptides are 25aa long, and in these cases gaps have been inserted at the front instead.
--MKQILLISLVVVLAVFAFNVAEG CDATC
KTHSVFGFFFKVLLIQVYLFNKTLA APSPI
---MARGAALALLLFGLLGVLVAAP DGGFD
MARMKYNIALIGILASVLLTIAVNA ENACN
-------MSFRSLLALSGLVCSGLA NVISK
--MSSTWIKFLFILTLVLLPYSVFS VNIFA
---------MIALFVLMGLMAAASA SSCCS
KKTVAALSFLFIVLFVAQEIAVTEA KTCEN
-----MKAIFVSALLVVALVASTSA HHQEL
------MLRLLLLPLFLFTLSMCMG QTFQY
|.... signal peptide ...|^next 5 amino acids - the peptide is cleaved at the "^" mark
Signal peptide comparison
OVERALL TASK:
Generating peptide LOGOs with the Logo Lab resource is as straightforward as before, and your goal is to compare the motifs in the signal peptides across the three taxonomical groups.
Play around with the options (dimensions etc) as much as you like, but make sure that:
It's easy to compare the plots (put in titles etc)
Plot the signal peptide part of the sequences in negative numbers
Read up on what the "Chemistry" coloring means (see below the plot) - this will help you interpret the results.
QUESTION 8:
Put the 3 plots in your report and comment on the following:
Which single positions (using the negative numbering scheme) appears to be the most important across all data sets?
Are there any regions where a certain class of amino acids appears to be important? What characterizes these amino acids?
Any striking differences between the three taxonomical groups?
Introducing Seq2Logo
2026 note: The new LogoLab webapp also provides the needed functionality for this part, but it's a little less tested on KL logos. Once you have done the part below with seq2logo (which we'll also be using while working with weight matrices) you're welcome to try to solve the same questions using LogoLab and give us feedback. Thanks.
As the next step we shall investigate how to work with the Seq2Logo tool hosted locally here at DTU. It offers a lot more advanced functionality with regards to the algorithms behind the plots, and many of these functions we'll get back to later in the course when we work with weight matrices.
TASK: create a basic LOGO plot for the eukaryote dataset
Set logo type to "Shannon"
Put in a title
Leave the rest of the "strange" options as default - for reasons we'll learn later they matter less for a large and well-balanced data set such as this one
QUESTION 9:
Include the plot in your report
Does it (more or less) show the same as the plot from LogoLab?
Small data sets
So far we have been working with relatively large data sets and have been able to see some quite clear patterns. But what happens if the data is limited? Below is a small set (n=20) of the signal peptides from the eukaryote data set: