You shall use bioinformatics tools to do a rational selection of peptides with a high potential as vaccine candidates against HCV in Guinea-Bissau. You shall tailor the selection towards the prevalent HLA-A, HLA-B, and HLA-DR molecules present in the population in Guinea-Bissau.
In detail you shall
HCV is a major cause of acute hepatitis and chronic liver disease, including cirrhosis and liver cancer. Globally, an estimated 170 million persons are chronically infected with HCV and 3 to 4 million persons are newly infected each year. HCV is spread primarily by direct contact with human blood. The major causes of HCV infection worldwide are use of unscreened blood transfusions, and re-use of needles and syringes that have not been adequately sterilized.
No vaccine is currently available to prevent hepatitis C and treatment for chronic hepatitis C is too costly for most persons in developing countries to afford. Thus, from a global perspective, the greatest impact on hepatitis C disease burden will likely be achieved by focusing efforts on reducing the risk of HCV transmission from nosocomial exposures (e.g. blood transfusions, unsafe injection practices) and high-risk behaviours (e.g. injection drug use).
Hepatitis C virus (HCV) is one of the viruses (A, B, C, D, and E), which together account for the vast majority of cases of viral hepatitis. It is an enveloped RNA virus in the flaviviridae family which appears to have a narrow host range. Humans and chimpanzees are the only known species susceptible to infection, with both species developing similar disease.
An important feature of the virus is the relative mutability of its genome, which in turn is probably related to the high propensity (80%) of inducing chronic infection. HCV is clustered into several distinct genotypes which may be important in determining the severity of the disease and the response to treatment.
For more details on HCV see link: WHO
Go to the Allele Frequency database
Select "HLA" tab, and next the "HLA classical allele freq search" submenu
Specify the HLA locus to A.
Set population Guinea-Bissau (Note, only Guinea-Bissau).
Set "Level of resolution" to >=2.
Set "Sort by" to Population, Highest to Lowest Frequency.
Leave other options to the default and press the search button.
Note, the HLA nomemclature is a little complicated. In short, the first two numbers after the locus name defines the protein name, and the remaining numbers different synonymous DNA substitutions with the coding region, and diffences in a non-coding region. That is two entries HLA-A*02:01:01 and HLA-A*02:01:02, encode an identical protein sequence but have one or more synonymous DNA substitutions with the coding region. For detail refer to HLA nomemclature.
Repeat this analysis for the HLA-B and HLA-DRB1 loci.
Go to GenBank.
Find the amino acid sequence for the YP_001469630.1 Hepatitis C virus genotype 2
poly-protein.
Save the sequence in
FASTA format. Note, make sure that your sequence is a protein sequence.
You shall use the NetMHCpan-4.1 prediction-server to identify potential CTL epitopes that will bind to your prevalent HLA-A and HLA-B molecules.
Go to the NetMHCpan-4.1 prediction-server and copy-paste the FASTA HCV genotype 2 genome protein sequence. Type in the HLA-A and HLA-B alleles you have found in question Q2 separated by commas (without blank spaces, and in the format HLA-A02:01!), select peptide length as 8,9, and 10mers, select Save prediction to XLS file, and press Submit. The calculation takes a while so please be patient.
Here is a link to an output file if NetMHCpan does not complete or if it takes too long for the calculation
to complete
NetMHCpan output file.
Once the calculation is completed, the output is displayed. Two prediction scores are provided for each peptide:MHC interaction (Score and %Rank). You can click on "Explain the output" on the lower part of the prediction output page for details on these scores. The peptide are classified as weak binders (WB) if the Rank is less than or equal to 2. Likewise are peptides with a Rank less than or equal to 0.5 classified as strong binders (SB). Try also to make sense of the other columns in the output
In the bottom of the results page you find a Link to output xls file. Open this file in excel.
In the file you will have two prediction score for each peptide for each allele (EL-score, and EL-Rank). The last column (NB) in the files contains the number of alleles to which the peptide-binding is classified as a weak or stronger binding (percentile rank score less than or equal to 2).
MS Excel notes:
Make sure that the decimal separator of your version of excel is set to '.' (periods).
Go to the NetMHCIIpan-4.0 prediction-server and upload the HCV genome sequence. Type in the HLA-DRB1 alleles you have found in question Q3 separated by commas (without blank spaces, and in the format DRB1_0301!), select Save prediction to xls file and press Submit. The calculation might takes some minutes.
Here is a link to the output if the server is down or the execution takes too long to complete NetMHCIIpan output
The NetMHCIIpan method provides two prediction scores for each peptide:MHC interaction Pred_Score and %Rank and the interpretation is the same as for the NetMHCpan methods. Here, however the peptide are classified as weak binders (WB) if the Rank is less than or equal to 5. Likewise are peptides with rank less than or equal to 1.0 classified as strong binders (SB).
In the bottom of the results page you find a Link to output xls file. Open this file in excel.
In the file, you have two prediction score for each peptide for each allele Score and Rank. The last column (NB) in the files contains the number of alleles to which the peptide-binding is classified as a weak or stronger binding.
You shall now investigate if any of you selected peptides are also present in an other HCV genotype. If this is the case, the peptides might not only have high importance for vaccine development against HCV in Guinea-Bissau, but also in other parts of the world.
Go to GenBank.
Find the amino acid sequence for the NC_004102 Hepatitis C virus genotype 1
poly-protein.
Save the sequence in
FASTA format. Note, make sure that your sequence is a protein sequence.
Hint: Copy-paste the translated sequence into an empty word document and remove all paragraphs and white-spaces. Then search for the 15mer sequence in the document.
As you have most likely realized duing this exercise, manually selecting epitope with a broad HLA coverarge is complicated. Also you have in this exercise demonstrated the importance of including sequence variation when identifying epitopes in order to ensure coverage of different strains of the given pathogen. A better approach is therefore to include all relevant HLA alleles and strain variations of the given pathogen in the search for potential epitopes, and next used methods such as PopCover to select peptide sets that in concert give optimal HLA allele and pathogenic strain coverage.
We will now work on applying the PopCover method to deal with these issues in an automated manner.
You shall use bioinformatics tools to do a rational selection of peptides with a high potential as vaccine candidates against HCV in Guinea-Bissau. You shall tailor the selection towards the prevalent HLA-A, HLA-B, and HLA-DR molecules present in the population in Guinea-Bissau. You shall do this using the PopCover method to identify broadly relevant T cell epitopes NetPolyEV-1.0 to construc an optimal polytope vaccine.
In detail you shall
In the first part of the exercise, you predicted HLA class I and class II epitope in the genome of HCV genotype 2. In case you didn't manage to complete this or have lost the prediction files, below are links to the output files
NetMHCpan output file (class I).
NetMHCIIpan output file (class II)
As described in the lecture, we have implemented a www-server implementation of the PopCover method at PopCover-2.0. Here, you can upload epitope data sets, and make peptide subset selection with optimal HLA and pathogen coverage. The input data to the server is a file in the format
Peptide HLA-allele [genetypeID]together with a file with HLA allele frequecy values.
If you are using a Linux/MAC PC, you can use the commands below to generate the epitope files. Simply save the raw output file from NetMHCpan, and run the commands
cat NetMHCpan-4.1.out | tr -d "*" | awk 'substr($2,1,3) == "HLA"' | awk '$13<=2' | awk '{print $3,$2,substr($11,1,12)}' > popcover_inp_classI
cat NetMHCIIpan-4.0.out | awk 'substr($2,1,4) == "DRB1"' | awk '$9<=5' | awk '{print $3,$2,substr($7,1,12),$5}' > popcover_inp_classII
where NetMHCpan-4.1.out and NetMHCIIpan-4.0.out are the raw output files produced by the NetMHCpan and NetMHCIIpan servers (i.e NOT the excel file).
Note, that we here apply the percentile rank threshold of 2.0% for class II and 5.0% for class II. These values can be altered if you what to be more stringent and
foces on the most likely binders only.
These commands will extract the peptide binders from the two prediction files, and print oput the predictions in a format suitable for the PopCover-2.0 server. Note that the third column in the two output files give the strain ID for the pathogen. This name has to be consistent between the two files, hence the substr($7,1,12) command.
If you work on a windows machine, you will not be able to run these commands. In this case, here are links to the generated data files
popcover_inp_classI file
popcover_inp_classII file
allele frequency file
Now use the PopCover-2.0 method,
to identify 5 peptides from Hepatitis C virus genotype 2 with optimal coverage of the HLA-A, B and DRB1
alleles you have selected. Use default options, except for "Use phenotypic frequencies for calculation" that should be selected.
Also, select Include protein sequence information (optional), and upload the HCV protein sequnece you have used in the first part of the exercise. Using this option, the PopCover method will map all the predicted epitope onto the set of 15mer (this is the default
size) peptides found in the source protein, and next apply these 15mer as input to the peptide search. Remember to ensure that the allele-names used for the HLA frequency information, have the same format as the allele-names in the epitope files.
If you have more time, run the predictions with NetMHCpan-4.1 for the HVC genotype 1 genome (for the
same HLA-I alleles as you used earlier), and use PopCover to identify a set of peptides with
optimal HLA-I AND pathogen coverage.
As the final step, you shall now contruct an optimal polytope. A polytope is a construct
where epitopes are placed "on a string" potentially including linker/spacer amino acids
to optimize construct (i.e minimize the number of introduced neo-epitopes).
You shall first try to construct such a polytope manually. Here you shall place the selected peptides from question Q12a, and add them into a protein sequence in a format
Next, you shall evaluate the properties of this initial construct in terms of the number of
predicted HLA-A class I neo-epitopes. We here only focus on a subset of the HLA-A alleles
in order to make the
computations run in a reasonable time. In an actual application, you should of course include
all the HLA-I alleles selected above.
Go to the
NetMHCpan-4.1 web-server. Upload your initial polytope construct, select your top 4 HLA-A alleles from question Q1 (the four HLA-A alleles with the highest frequencies), set Filtering threshold for %Rank (leave -99 to print all) to
2 (to limit the output produced), and submit the prediction.
Investigate output.
If this is the case, go back and modify the polytope by reordering the peptide order and/or introducing some short linkers between the peptides, i.e.
You have now seen that this manual approach is highly troublesome and time consuming.
To resolve this, we have developed a Polytope Optimzer called NetPolyEV. The
tool is implemented as a web-server available at NetPolyEV-1.0.
The tool is still in development and is in many aspect not in an optimal state. For
instance is the run-time very long. However, we will here try to use the tool to give you
an idea of its potential.
Go to the NetPolyEV-1.0 website. Select the Load sample data option, select the two alleles HLA-A*02:01 and HLA-A*24:02. Leave all other options unchanged, and submit. Note, the the submission is slow and will take ~1 min to complete. Look at the result.
Now select Monte Carlo iterations at each temperature: equal to 5, and press Improve with slow Monte Carlo method. This will launch a slow and more time consuming
polytope optimizier. Again, please be patient, this run will take 2-4 mins to complete. Look at the result.
Finally, try to submit the 5 peptides selected in Q12, and run the NetPolyEV with
the top 4 HLA-A alleles identified in Q1.
Now you are done!! You can try to sell your peptides to big pharma and get rich.
>MY_INITIAL_POLYTOPE
pep1
pep2
pep3
...
>MY_SECOND_POLYTOPE
pep3
lnk1
pep1
lnk2
pep2
...
where lnk is a short linker of amino acids of length 1-6. You can try this manual
optimization a few time.
If you are interested in running the NetPolyEV optimizer from the commandline, a stand-alone
version is available from GitHub - NetPolyEV.