Exercise:AntigenProcessing ans
Answers
Get the Data and Filter It
- Q1: How many epitopes do you find with the initial filter?
Answer: 646 Epitopes.
- Q2: Do you agree with calling these peptides epitopes? Why?
Answer: Not really. It would be more precise to call them MHC ligands or HLA-presented peptides.
Explanation: Mass spectrometry demonstrates that a peptide was isolated in association with an MHC molecule and therefore provides evidence that the peptide is naturally processed and presented. However, this does not by itself demonstrate that the peptide is recognized by a T-cell receptor or induces a T-cell response.
An epitope generally refers to a molecular structure that is recognized by the adaptive immune system. Therefore, an MHC ligand only becomes a demonstrated T-cell epitope when there is evidence of T-cell recognition.
- Q3: How many eluted ligands did they find in this study?
Answer: Wang et al. found 2274 HLA-DR presented peptides
Be aware that the number of non-redundant presented peptides reported in the paper is not be identical to the IEDB number: 1,593. Why is that?
- Q4: What kind of post-translational modifications are present?
Answer: The dataset contains modified peptide residues rather than ordinary unmodified peptide sequences. One relevant modification in these RA HLA-DR datasets is cysteinylation, in which a cysteine residue forms a disulfide with a free cysteine.
Explanation: A post-translational modification changes the chemical composition and mass of a peptide. This is important for motif deconvolution because the prediction models expect ordinary amino-acid sequences and generally do not represent modified residues in the same way.
The underlying PXD003051 dataset includes searches for modifications such as cysteinylation, deamidation, hydroxylation and other modified residues. For the exercise, students should report the exact modification labels shown in the Modifications column of their IEDB export.
The modified peptides are removed before continuing because MHCMotifDecon expects standard peptide sequences.
- Q5: Which alleles are expressed in the patient used for this assay?
Answer: For the RA sample used in this part of the exercise, the relevant HLA-DRB1 genotype is:
HLA-DRB1*01:01 HLA-DRB1*04:01
For MHCMotifDecon these are entered as:
DRB1_0101 DRB1_0401
Explanation: Humans are normally heterozygous at HLA loci, so a patient can express more than one HLA-DRB1 molecule. Consequently, an immunopeptidomics experiment performed on material from that patient contains a mixture of peptides presented by the different HLA molecules.
This is exactly why motif deconvolution is required: from the mass-spectrometry experiment alone, we know which peptides were isolated, but we do not initially know which of the co-expressed HLA molecules presented each individual peptide.
MHCMotifDecon
- Q6: How many peptides were assigned to each of the alleles, and how many were assigned to the trash cluster?
Answer: Record the peptide numbers shown in the MHCMotifDecon output for:
DRB1_0101 = [number from your run] DRB1_0401 = [number from your run] Trash = [number from your run]
Explanation: These numbers should be taken directly from the analysis rather than hard-coded because they depend on the exact input peptide list, filtering and chosen %Rank threshold.
MHCMotifDecon predicts the binding of every peptide to each supplied HLA molecule and assigns the peptide to its most likely restriction element. With the default settings, peptides that have a predicted %Rank greater than 20 for all supplied HLA molecules are assigned to the trash cluster.
Therefore, the trash cluster is not a third HLA allele. It represents peptides for which none of the supplied HLA molecules provides a sufficiently convincing predicted binding interaction.
- Q7: Can you see any differences between the trash-cluster peptides and the peptides assigned to a specific allele? Explain.
Answer: Yes. Peptides assigned to a particular HLA allele usually form a more recognizable HLA-binding motif, whereas the trash peptides tend to show a much weaker or less coherent sequence pattern.
Explanation: HLA-DR molecules bind peptides through a characteristic peptide-binding core. Certain amino acids are preferred at particular anchor positions, especially within the approximately nine-residue binding core. When many peptides presented by the same HLA molecule are aligned, these preferences produce a recognizable sequence logo.
The trash cluster contains peptides that are poorly predicted to bind any of the HLA molecules provided to MHCMotifDecon. Consequently, they often:
- lack a clear HLA-binding motif;
- have weaker predicted binding scores;
- contain more heterogeneous sequences; and
- may include co-immunoprecipitated contaminants or peptides presented by an HLA molecule that was not included in the analysis.
The trash cluster is therefore useful: instead of forcing every peptide into one of the specified HLA motifs, MHCMotifDecon can separate sequences that do not fit the proposed HLA repertoire.
- Q8: What is the effect of changing the %Rank threshold?
Answer: Changing the %Rank threshold changes how stringent MHCMotifDecon is when deciding whether a peptide should be assigned to an HLA molecule or placed in the trash cluster.
Explanation:
A lower %Rank threshold is more stringent:
- fewer peptides are accepted as HLA binders;
- more peptides enter the trash cluster; and
- the resulting HLA motifs are generally cleaner, but some real ligands may be discarded.
A higher %Rank threshold is less stringent:
- more peptides are assigned to HLA alleles;
- fewer peptides enter the trash cluster; and
- weaker predicted binders and potentially contaminating peptides may be included.
Remember that for %Rank, smaller values indicate stronger predicted binding. The default MHCMotifDecon trash threshold for MHC class II is 20%.
The exercise therefore illustrates the trade-off between sensitivity and specificity: very stringent filtering produces cleaner assignments but risks losing true ligands, whereas permissive filtering retains more peptides but may introduce noise.
- Q9: How many peptides are now associated with HLA-DRB4? Where did they come from?
Answer: The exact DRB4 peptide count should be recorded from the second MHCMotifDecon run:
DRB4_0101 = [number from your run]
The important observation is that adding DRB4_0101 causes a subset of peptides to be reassigned to DRB4.
Where did they come from? They generally come from peptides that, in the first analysis, had been:
- assigned to one of the DRB1 molecules, particularly because DRB4 was not available as an alternative; or
- placed in the trash cluster because neither DRB1*01:01 nor DRB1*04:01 provided a sufficiently good prediction.
Explanation: HLA-DRB1*04:01 is commonly linked with expression of HLA-DRB4*01:01. If DRB4 is genuinely expressed but is omitted from the HLA typing supplied to the algorithm, MHCMotifDecon has no opportunity to assign peptides to it.
This demonstrates an important principle of supervised motif deconvolution:
The quality of the deconvolution depends on the completeness of the HLA typing supplied to the method.
Once DRB4 is included, peptides fitting its binding specificity can form a separate motif rather than being incorrectly associated with another allele or classified as trash.
DRB4 normally contributes a smaller peptide repertoire than the accompanying major DRB1 molecule, which also helps explain why its motif can be harder to identify with an unsupervised method.
GibbsCluster
- Q10: Does the number of clusters found by GibbsCluster correspond to the number of HLA alleles in your sample?
Answer: Not necessarily. The number of sequence clusters identified by GibbsCluster does not have to equal the number of HLA molecules expressed by the sample.
In this experiment, after taking DRB4 into account, there are three relevant HLA-DR specificities:
DRB1*01:01 DRB1*04:01 DRB4*01:01
However, an unsupervised clustering analysis may recover only the strongest distinguishable motifs and may fail to resolve a smaller peptide population such as the DRB4 repertoire as an independent cluster.
Explanation: GibbsCluster does not know the patient's HLA genotype. It only examines patterns in the peptide sequences and asks how the sequences can best be partitioned into clusters.
Therefore:
number of sequence clusters ≠ necessarily number of HLA alleles
Two alleles with similar motifs may be merged into one cluster, a small allele-specific repertoire may not form a sufficiently strong independent cluster, or noisy peptides may affect the clustering solution.
This contrasts with MHCMotifDecon, where the known HLA molecules are explicitly supplied to the algorithm.
- Q11: Can you identify other differences between the GibbsCluster solution and the MHCMotifDecon solution?
Answer: Yes. The most important difference is that MHCMotifDecon is supervised by HLA-binding predictions and known HLA typing, whereas GibbsCluster is an unsupervised sequence-clustering method.
Explanation:
With MHCMotifDecon:
- the HLA alleles expressed by the sample are supplied beforehand;
- each peptide is evaluated using allele-specific binding predictions;
- clusters are directly labelled with specific HLA molecules;
- peptides that do not bind any supplied allele sufficiently well can be placed in a trash cluster; and
- relatively small allele-specific peptide populations may still be detected because the algorithm already knows which alleles to test.
With GibbsCluster:
- no HLA genotype information is required;
- peptides are grouped according to similarities in their sequence motifs;
- the resulting clusters are not intrinsically labelled as particular HLA alleles;
- allele identities have to be inferred afterwards by comparing the motifs with known HLA-binding motifs; and
- weak or small motifs can be merged with larger clusters or may not emerge as independent clusters.
Therefore, the two approaches answer slightly different questions.
MHCMotifDecon asks: Given the HLA molecules that I know are present, which HLA molecule most likely presented each peptide?
GibbsCluster asks: Without knowing which HLA molecules are present, how many different sequence patterns can I detect in this peptide dataset?
Using both approaches is useful because agreement between them gives additional confidence in the inferred motifs, while disagreement can reveal incomplete HLA typing, low-abundance HLA molecules, overlapping binding specificities, or contaminants.