HCV vaccine development - Part 1

Rational Vaccine design

You shall use bioinformatics tools to do a rational selection of peptides with a high potential as vaccine candidates against HCV in Guinea-Bissau. You shall tailor the selection towards the prevalent HLA-A, HLA-B, and HLA-DR molecules present in the population in Guinea-Bissau.

In detail you shall

  1. Identification of the prevalent HLA-A, HLA-B and HLA-DR alleles in Guinea-Bissau.
  2. Down-load the prevalent HCV genotype in Africa.
  3. Select a minimal set of peptides tailored to bind to the prevalent HLA types.
  4. Check if the selected peptides are cross-reactive towards other HCV subtypes.

Background

HCV is a major cause of acute hepatitis and chronic liver disease, including cirrhosis and liver cancer. Globally, an estimated 170 million persons are chronically infected with HCV and 3 to 4 million persons are newly infected each year. HCV is spread primarily by direct contact with human blood. The major causes of HCV infection worldwide are use of unscreened blood transfusions, and re-use of needles and syringes that have not been adequately sterilized.

No vaccine is currently available to prevent hepatitis C and treatment for chronic hepatitis C is too costly for most persons in developing countries to afford. Thus, from a global perspective, the greatest impact on hepatitis C disease burden will likely be achieved by focusing efforts on reducing the risk of HCV transmission from nosocomial exposures (e.g. blood transfusions, unsafe injection practices) and high-risk behaviours (e.g. injection drug use).

Hepatitis C virus (HCV) is one of the viruses (A, B, C, D, and E), which together account for the vast majority of cases of viral hepatitis. It is an enveloped RNA virus in the flaviviridae family which appears to have a narrow host range. Humans and chimpanzees are the only known species susceptible to infection, with both species developing similar disease.

An important feature of the virus is the relative mutability of its genome, which in turn is probably related to the high propensity (80%) of inducing chronic infection. HCV is clustered into several distinct genotypes which may be important in determining the severity of the disease and the response to treatment.

For more details on HCV see link: WHO


The exercise

Identification of prevalent HLA-A, HLA-B, and HLA-DR alleles in Guinea-Bissau.

Go to the Allele Frequency database

Select "HLA" tab, and next the "HLA classical allele freq search" submenu

Specify the HLA locus to A.
Set population Guinea-Bissau (Note, only Guinea-Bissau).

Set "Level of resolution" to >=2.
Set "Sort by" to Population, Highest to Lowest Frequency.

Leave other options to the default and press the search button.

Note, the HLA nomemclature is a little complicated. In short, the first two numbers after the locus name defines the protein name, and the remaining numbers different synonymous DNA substitutions with the coding region, and diffences in a non-coding region. That is two entries HLA-A*02:01:01 and HLA-A*02:01:02, encode an identical protein sequence but have one or more synonymous DNA substitutions with the coding region. For detail refer to HLA nomemclature.

  • Q1 Which are the most prevalent HLA-A alleles in Guinea-Bissau (allele frequency >= 5%)? Note that you only need to report the first four digits of the HLA typing, i.e A*02:01.

    Repeat this analysis for the HLA-B and HLA-DRB1 loci.

  • Q2 Which are the most prevalent (allele frequency >= 5%) HLA-B alleles in Guinea-Bissau?
  • Q3 Which are the most prevalent (allele frequency >= 5%) HLA-DRB1 alleles in Guinea-Bissau?

    Download the prevalent HCV genotype in Africa.

    You now need to identify a representative genomic sequence for the most abundant HCV genotype in the selected population. In this case it is genotype 2. We need the sequence in protein fasta format.

    Go to GenBank.
    Find the amino acid sequence for the YP_001469630.1 Hepatitis C virus genotype 2 poly-protein.
    Save the sequence in FASTA format. Note, make sure that your sequence is a protein sequence.


    Identification of CTL epitopes

    You shall use the NetMHCpan-4.1 prediction-server to identify potential CTL epitopes that will bind to your prevalent HLA-A and HLA-B molecules.

    Go to the NetMHCpan-4.1 prediction-server and copy-paste the FASTA HCV genotype 2 genome protein sequence. Type in the HLA-A and HLA-B alleles you have found in question Q2 separated by commas (without blank spaces, and in the format HLA-A02:01!), select peptide length as 8,9, and 10mers, select Save prediction to XLS file, and press Submit. The calculation takes a while so please be patient.

    Here is a link to an output file if NetMHCpan does not complete or if it takes too long for the calculation to complete NetMHCpan output file.

    Once the calculation is completed, the output is displayed. Two prediction scores are provided for each peptide:MHC interaction (Score and %Rank). You can click on "Explain the output" on the lower part of the prediction output page for details on these scores. The peptide are classified as weak binders (WB) if the Rank is less than or equal to 2. Likewise are peptides with a Rank less than or equal to 0.5 classified as strong binders (SB). Try also to make sense of the other columns in the output

    In the bottom of the results page you find a Link to output xls file. Open this file in excel.

    In the file you will have two prediction score for each peptide for each allele (EL-score, and EL-Rank). The last column (NB) in the files contains the number of alleles to which the peptide-binding is classified as a weak or stronger binding (percentile rank score less than or equal to 2).

    MS Excel notes:
    Make sure that the decimal separator of your version of excel is set to '.' (periods).

  • Q5a Do you observe length differences in the high scoring peptides from the different alleles? I.e. how many peptides with length different from 9 do you find within the top 20 peptides for each allele? Hint, add a column with the peptide length, sort next for each allele in the EL_Rank score from low to high, and report the numver of 8mer and 10mer within the top 20 peptides for each allele.
  • Q5b Identify 5 peptides that will give you the broadest allelic coverage.
  • Q6 How many of the five peptides binder to each of the HLA-A and B alleles? (Help: construct a table with columns corresponding to each HLA and rows corresponding to each peptide, and make crosses in cells corresponding to peptide:HLA binders)

    Identification of T-helper epitopes

    To identify potential T-helper epitopes you shall use the NetMHCIIpan-4.0 prediction-server. Note, as described in the lectures, a more recent version of this tool (version 4.3) exists. We will here, for sake of consistency with the results following later, use the older version. But for any future work, you are encouraged to use the most recent version.

    Go to the NetMHCIIpan-4.0 prediction-server and upload the HCV genome sequence. Type in the HLA-DRB1 alleles you have found in question Q3 separated by commas (without blank spaces, and in the format DRB1_0301!), select Save prediction to xls file and press Submit. The calculation might takes some minutes.

    Here is a link to the output if the server is down or the execution takes too long to complete NetMHCIIpan output

    The NetMHCIIpan method provides two prediction scores for each peptide:MHC interaction Pred_Score and %Rank and the interpretation is the same as for the NetMHCpan methods. Here, however the peptide are classified as weak binders (WB) if the Rank is less than or equal to 5. Likewise are peptides with rank less than or equal to 1.0 classified as strong binders (SB).

    In the bottom of the results page you find a Link to output xls file. Open this file in excel.

    In the file, you have two prediction score for each peptide for each allele Score and Rank. The last column (NB) in the files contains the number of alleles to which the peptide-binding is classified as a weak or stronger binding.

  • Q7 Are any of the CTL epitopes you found in question Q5 part of a 15mer T-helper epitope for any of the HLA-DRB1 alleles?

  • Q8 Select the smallest number of 15mer peptides so that these peptides will have coverage of all HLA-A, HLA-B, and HLA-DRB1 alleles in your prevalent selection

    Known CTL epitopes

    Now you should go into the lab and investigate if any of the predicted epitopes are indeed CTL and/or T-helper epitopes. However before doing this, it might be informative to check the public epitope database for information on whether other groups have analyzed the peptides. Go to the Immune Epitope database (IEDB). Upload each of the selected 15mer peptides (as "Linear Peptide:"), select "Substring" similarity.

  • Q9 Are any of the 15mer peptides described in the IEDB? And if yes does the restriction element (i.e. the HLA) and epitope match the predictions?

    Peptide conservation in other HCV subtypes

    You shall now investigate if any of you selected peptides are also present in an other HCV genotype. If this is the case, the peptides might not only have high importance for vaccine development against HCV in Guinea-Bissau, but also in other parts of the world.

    Go to GenBank.
    Find the amino acid sequence for the NC_004102 Hepatitis C virus genotype 1 poly-protein.
    Save the sequence in FASTA format. Note, make sure that your sequence is a protein sequence.

  • Q10 Are any of the 15mers selected in question Q8 present in the HVC genotype 1 genome?
    Hint: Copy-paste the translated sequence into an empty word document and remove all paragraphs and white-spaces. Then search for the 15mer sequence in the document.
  • Q11 What does this imply for the worldwide coverage of your vaccine?

    As you have most likely realized duing this exercise, manually selecting epitope with a broad HLA coverarge is complicated. Also you have in this exercise demonstrated the importance of including sequence variation when identifying epitopes in order to ensure coverage of different strains of the given pathogen. A better approach is therefore to include all relevant HLA alleles and strain variations of the given pathogen in the search for potential epitopes, and next used methods such as PopCover to select peptide sets that in concert give optimal HLA allele and pathogenic strain coverage.

    We will now work on applying the PopCover method to deal with these issues in an automated manner.


    HCV vaccine development - Part 2. PopCover and Polytope-optimizer

    Rational Vaccine design

    You shall use bioinformatics tools to do a rational selection of peptides with a high potential as vaccine candidates against HCV in Guinea-Bissau. You shall tailor the selection towards the prevalent HLA-A, HLA-B, and HLA-DR molecules present in the population in Guinea-Bissau. You shall do this using the PopCover method to identify broadly relevant T cell epitopes NetPolyEV-1.0 to construc an optimal polytope vaccine.

    In detail you shall

    1. Identify 15mer peptides with optimal HLA-A, HLA-B and HLA-DR allelic coverage
    2. Construct a polytope with optimal antigenic behavior

    Background

    In the first part of the exercise, you predicted HLA class I and class II epitope in the genome of HCV genotype 2. In case you didn't manage to complete this or have lost the prediction files, below are links to the output files

    NetMHCpan output file (class I).
    NetMHCIIpan output file (class II)

    As described in the lecture, we have implemented a www-server implementation of the PopCover method at PopCover-2.0. Here, you can upload epitope data sets, and make peptide subset selection with optimal HLA and pathogen coverage. The input data to the server is a file in the format

    Peptide	HLA-allele [genetypeID]
    
    together with a file with HLA allele frequecy values.

    If you are using a Linux/MAC PC, you can use the commands below to generate the epitope files. Simply save the raw output file from NetMHCpan, and run the commands

    cat NetMHCpan-4.1.out | tr -d "*" | awk 'substr($2,1,3) == "HLA"' | awk '$13<=2' | awk '{print $3,$2,substr($11,1,12)}' > popcover_inp_classI
    cat NetMHCIIpan-4.0.out | awk 'substr($2,1,4) == "DRB1"' | awk '$9<=5' | awk '{print $3,$2,substr($7,1,12),$5}' > popcover_inp_classII
    
    where NetMHCpan-4.1.out and NetMHCIIpan-4.0.out are the raw output files produced by the NetMHCpan and NetMHCIIpan servers (i.e NOT the excel file). Note, that we here apply the percentile rank threshold of 2.0% for class II and 5.0% for class II. These values can be altered if you what to be more stringent and foces on the most likely binders only.

    These commands will extract the peptide binders from the two prediction files, and print oput the predictions in a format suitable for the PopCover-2.0 server. Note that the third column in the two output files give the strain ID for the pathogen. This name has to be consistent between the two files, hence the substr($7,1,12) command.

    If you work on a windows machine, you will not be able to run these commands. In this case, here are links to the generated data files

    popcover_inp_classI file
    popcover_inp_classII file
    allele frequency file

    Now use the PopCover-2.0 method, to identify 5 peptides from Hepatitis C virus genotype 2 with optimal coverage of the HLA-A, B and DRB1 alleles you have selected. Use default options, except for "Use phenotypic frequencies for calculation" that should be selected. Also, select Include protein sequence information (optional), and upload the HCV protein sequnece you have used in the first part of the exercise. Using this option, the PopCover method will map all the predicted epitope onto the set of 15mer (this is the default size) peptides found in the source protein, and next apply these 15mer as input to the peptide search. Remember to ensure that the allele-names used for the HLA frequency information, have the same format as the allele-names in the epitope files.

  • Q12a What peptides did PopCover select, and how do the selected peptides differ from the ones you selected manually?
  • Q12b That is, what how many times are each of the HLA alleles targeted by your and by the PopCover selection, respectively?
  • If you have more time, run the predictions with NetMHCpan-4.1 for the HVC genotype 1 genome (for the same HLA-I alleles as you used earlier), and use PopCover to identify a set of peptides with optimal HLA-I AND pathogen coverage.


    As the final step, you shall now contruct an optimal polytope. A polytope is a construct where epitopes are placed "on a string" potentially including linker/spacer amino acids to optimize construct (i.e minimize the number of introduced neo-epitopes).

    You shall first try to construct such a polytope manually. Here you shall place the selected peptides from question Q12a, and add them into a protein sequence in a format

    >MY_INITIAL_POLYTOPE
    pep1
    pep2
    pep3
    ...
    

    Next, you shall evaluate the properties of this initial construct in terms of the number of predicted HLA-A class I neo-epitopes. We here only focus on a subset of the HLA-A alleles in order to make the computations run in a reasonable time. In an actual application, you should of course include all the HLA-I alleles selected above.

    Go to the NetMHCpan-4.1 web-server. Upload your initial polytope construct, select your top 4 HLA-A alleles from question Q1 (the four HLA-A alleles with the highest frequencies), set Filtering threshold for %Rank (leave -99 to print all) to 2 (to limit the output produced), and submit the prediction.

    Investigate output.

  • Q13 Does your initial polytope construct contain any neo-epitopes (i.e. peptides not in the initial epitope selected predicted bind one or more of the included HLAs)?

    If this is the case, go back and modify the polytope by reordering the peptide order and/or introducing some short linkers between the peptides, i.e.

    >MY_SECOND_POLYTOPE
    pep3
    lnk1
    pep1
    lnk2
    pep2
    ...
    
    where lnk is a short linker of amino acids of length 1-6. You can try this manual optimization a few time.

  • Q14 Were you able to identify a polytope construct without any predicted neo-epitopes?

    You have now seen that this manual approach is highly troublesome and time consuming. To resolve this, we have developed a Polytope Optimzer called NetPolyEV. The tool is implemented as a web-server available at NetPolyEV-1.0.

    The tool is still in development and is in many aspect not in an optimal state. For instance is the run-time very long. However, we will here try to use the tool to give you an idea of its potential.

    Go to the NetPolyEV-1.0 website. Select the Load sample data option, select the two alleles HLA-A*02:01 and HLA-A*24:02. Leave all other options unchanged, and submit. Note, the the submission is slow and will take ~1 min to complete. Look at the result.

  • Q15 How many neo-epitopes was found in the original epitope, and how many in the refined

    Now select Monte Carlo iterations at each temperature: equal to 5, and press Improve with slow Monte Carlo method. This will launch a slow and more time consuming polytope optimizier. Again, please be patient, this run will take 2-4 mins to complete. Look at the result.

  • Q16 How many neo-epitopes was found in the final epitope?

    Finally, try to submit the 5 peptides selected in Q12, and run the NetPolyEV with the top 4 HLA-A alleles identified in Q1.

  • Q17 Was NetPolyEV better than you in identifying an optimal polytope construct (that is was the number of neo-epitopes in the NetPolyEV polytope construct lower that what you could find in question?
    If you are interested in running the NetPolyEV optimizer from the commandline, a stand-alone version is available from GitHub - NetPolyEV.

    Now you are done!! You can try to sell your peptides to big pharma and get rich.