Exercise:AntigenProcessing ans: Difference between revisions

From 22145
Jump to navigation Jump to search
No edit summary
 
(22 intermediate revisions by the same user not shown)
Line 5: Line 5:
*'''Q1: How many epitopes do you find with the initial filter?'''
*'''Q1: How many epitopes do you find with the initial filter?'''


'''Answer:''' The exact number should be recorded from the IEDB search at the time the exercise is performed.
'''Answer:''' 646 Epitopes.
 
This number should not be hard-coded into the exercise because the IEDB is continuously updated and re-curated. The important point is that the initial search retrieves MHC ligand records associated with rheumatoid arthritis in humans, before applying the more restrictive mass-spectrometry and HLA-DR filters.
 
'''Explanation:''' At this stage the search is deliberately broad. Subsequent filters progressively reduce the dataset to the particular immunopeptidomics experiment that we want to analyse.


*'''Q2: Do you agree with calling these peptides epitopes? Why?'''
*'''Q2: Do you agree with calling these peptides epitopes? Why?'''


'''Answer:''' Not necessarily. It would be more precise to call them '''MHC ligands''' or '''HLA-presented peptides'''.
'''Answer:''' Not really. It would be more precise to call them '''MHC ligands''' or '''HLA-presented peptides'''.


'''Explanation:''' Mass spectrometry demonstrates that a peptide was isolated in association with an MHC molecule and therefore provides evidence that the peptide is naturally processed and presented. However, this does not by itself demonstrate that the peptide is recognized by a T-cell receptor or induces a T-cell response.
'''Explanation:''' Mass spectrometry demonstrates that a peptide was isolated in association with an MHC molecule and therefore provides evidence that the peptide is naturally processed and presented. However, this does not by itself demonstrate that the peptide is recognized by a T-cell receptor or induces a T-cell response.


An '''epitope''' generally refers to a molecular structure that is recognized by the adaptive immune system. Therefore, an MHC ligand only becomes a demonstrated T-cell epitope when there is evidence of T-cell recognition.
An '''epitope''' generally refers to a molecular structure that is recognized by the adaptive immune system. Therefore, an MHC ligand only becomes a demonstrated T-cell epitope when there is evidence of T-cell recognition.
This distinction is particularly clear in the Wang et al. study: many HLA-DR-presented peptides were identified by mass spectrometry, but only a subset was subsequently shown to be immunogenic.


*'''Q3: How many eluted ligands did they find in this study?'''
*'''Q3: How many eluted ligands did they find in this study?'''


'''Answer:''' Wang et al. reported '''1,593 non-redundant HLA-DR-presented peptides''', originating from '''870 source proteins'''.
'''Answer:''' Wang et al. found  2274 HLA-DR presented peptides


'''Explanation:''' These peptides were identified by LC-MS/MS from synovial tissue, synovial fluid mononuclear cells, and peripheral blood mononuclear cells obtained from patients with rheumatoid arthritis and Lyme arthritis.
Be aware that the number of non-redundant presented peptides reported in the paper is not be identical to the IEDB number: 1,593. Why is that?
 
Be aware that the number displayed in a particular IEDB results table may not be identical to 1,593. IEDB may represent individual assays, peptide records, or a filtered subset of the publication, whereas 1,593 is the total number of non-redundant HLA-DR-presented peptides reported in the publication.


*'''Q4: What kind of post-translational modifications are present?'''
*'''Q4: What kind of post-translational modifications are present?'''


'''Answer:''' The dataset contains '''modified peptide residues rather than ordinary unmodified peptide sequences'''. One relevant modification in these RA HLA-DR datasets is '''cysteinylation''', in which a cysteine residue forms a disulfide with a free cysteine.
+ DEAM(N2) and + CITR(R8) Deamination and Citrulation


'''Explanation:''' A post-translational modification changes the chemical composition and mass of a peptide. This is important for motif deconvolution because the prediction models expect ordinary amino-acid sequences and generally do not represent modified residues in the same way.


The underlying PXD003051 dataset includes searches for modifications such as cysteinylation, deamidation, hydroxylation and other modified residues. For the exercise, students should report the exact modification labels shown in the '''Modifications''' column of their IEDB export.


The modified peptides are removed before continuing because MHCMotifDecon expects standard peptide sequences.
*'''Q5: Which alleles are expressed in the patient used for this assay?'''


*'''Q5: Which alleles are expressed in the patient used for this assay?'''
'''Answer:''' Each RA sample has different alleles:


'''Answer:''' For the RA sample used in this part of the exercise, the relevant HLA-DRB1 genotype is:
RA1    DRB1_0402,DRB1_1104
RA2    DRB1_0101,DRB1_0401
RA3    DRB1_0101,DRB1_0401
RA4    DRB1_0401,DRB1_1501
RA5    DRB1_0801,DRB1_1501


HLA-DRB1*01:01
=== MHCMotifDecon ===
HLA-DRB1*04:01


For MHCMotifDecon these are entered as:
*'''Q6: How many peptides were assigned to each of the alleles, and how many were assigned to the trash cluster?'''


DRB1_0101
[[Image:MHCMotifDecon_rank20.png|center|800px]]
DRB1_0401


'''Explanation:''' Humans are normally heterozygous at HLA loci, so a patient can express more than one HLA-DRB1 molecule. Consequently, an immunopeptidomics experiment performed on material from that patient contains a mixture of peptides presented by the different HLA molecules.


This is exactly why motif deconvolution is required: from the mass-spectrometry experiment alone, we know which peptides were isolated, but we do not initially know which of the co-expressed HLA molecules presented each individual peptide.
RA1 DRB1_0402=141,DRB1_1104=71 Trash=18
RA2 DRB1_0101=193,DRB1_0401=245 Trash=52
RA3 DRB1_0101=124,DRB1_0401=301 Trash=11
RA4 DRB1_0401=51,DRB1_1501=16 Trash=3
RA5 DRB1_0801=101,DRB1_1501=31 Trash=7


=== MHCMotifDecon ===


*'''Q6: How many peptides were assigned to each of the alleles, and how many were assigned to the trash cluster?'''
A total of '''1365 peptides''' were analysed. Although 1496 sequences were submitted, only peptides within the selected length range of '''12–21 amino acids''' were included in the deconvolution.


'''Answer:''' Record the peptide numbers shown in the MHCMotifDecon output for:
*'''Q7: Can you see any differences between the trash-cluster peptides and the peptides assigned to a specific allele? Explain.'''


DRB1_0101 = [number from your run]
'''Answer:''' Yes. Peptides assigned to specific HLA alleles show clear and reproducible '''binding motifs''', with strong amino-acid preferences at particular positions of the peptide-binding core.
DRB1_0401 = [number from your run]
Trash      = [number from your run]


'''Explanation:''' These numbers should be taken directly from the analysis rather than hard-coded because they depend on the exact input peptide list, filtering and chosen %Rank threshold.
In contrast, the Trash clusters show much weaker and less defined sequence motifs. These peptides are predicted to bind poorly to all HLA alleles included for that patient and may therefore represent weak binders, contaminants, or peptides presented by an HLA molecule that was not included in the analysis.


MHCMotifDecon predicts the binding of every peptide to each supplied HLA molecule and assigns the peptide to its most likely restriction element. With the default settings, peptides that have a predicted '''%Rank greater than 20 for all supplied HLA molecules''' are assigned to the '''trash cluster'''.
For RA4 and RA5, only 3 and 7 peptides, respectively, were assigned to Trash. Because MHCMotifDecon requires at least 10 sequences to generate a logo, no Trash logo was produced for these patients.


Therefore, the trash cluster is not a third HLA allele. It represents peptides for which none of the supplied HLA molecules provides a sufficiently convincing predicted binding interaction.
*'''Q8: What is the effect of changing the %Rank threshold?'''


*'''Q7: Can you see any differences between the trash-cluster peptides and the peptides assigned to a specific allele? Explain.'''
'''Answer:''' Lowering the %Rank threshold from '''20% to 2%''' makes the peptide assignment considerably more stringent. Peptides with a best predicted binding rank between 2% and 20% are no longer accepted as binders and are instead assigned to the '''Trash''' cluster.


'''Answer:''' Yes. Peptides assigned to a particular HLA allele usually form a more recognizable '''HLA-binding motif''', whereas the trash peptides tend to show a much weaker or less coherent sequence pattern.
[[Image:MHCMotifDecon_rank2.png|center|800px]]


'''Explanation:''' HLA-DR molecules bind peptides through a characteristic peptide-binding core. Certain amino acids are preferred at particular anchor positions, especially within the approximately nine-residue binding core. When many peptides presented by the same HLA molecule are aligned, these preferences produce a recognizable sequence logo.
The effect can be seen directly by comparing the two runs:


The trash cluster contains peptides that are poorly predicted to bind any of the HLA molecules provided to MHCMotifDecon. Consequently, they often:
RA1 Trash: 18 → 51 (+33)
RA2 Trash: 52 → 91 (+39)
RA3 Trash: 11 → 34 (+23)
RA4 Trash: 3 → 11 (+8)
RA5 Trash: 7 → 37 (+30)


* lack a clear HLA-binding motif;
Across all five patients:
* have weaker predicted binding scores;
* contain more heterogeneous sequences; and
* may include co-immunoprecipitated contaminants or peptides presented by an HLA molecule that was not included in the analysis.


The trash cluster is therefore useful: instead of forcing every peptide into one of the specified HLA motifs, MHCMotifDecon can separate sequences that do not fit the proposed HLA repertoire.
%Rank 20 91 peptides in Trash
%Rank 2 224 peptides in Trash


*'''Q8: What is the effect of changing the %Rank threshold?'''
Thus, lowering the threshold to 2% moved '''133 additional peptides''' into the Trash cluster.


'''Answer:''' Changing the %Rank threshold changes how stringent MHCMotifDecon is when deciding whether a peptide should be assigned to an HLA molecule or placed in the trash cluster.
Correspondingly, the number of peptides assigned to the HLA-DRB1 alleles decreased:


'''Explanation:'''
RA1 DRB1_0402: 141 → 120 DRB1_1104: 71 → 59
RA2 DRB1_0101: 193 → 177 DRB1_0401: 245 → 222
RA3 DRB1_0101: 124 → 118 DRB1_0401: 301 → 284
RA4 DRB1_0401: 51 → 44 DRB1_1501: 16 → 15
RA5 DRB1_0801: 101 → 77 DRB1_1501: 31 → 25


A '''lower %Rank threshold''' is more stringent:
The total number of peptides assigned to an HLA allele therefore decreased from '''1274 at the 20% threshold''' to '''1141 at the 2% threshold'''.


* fewer peptides are accepted as HLA binders;
Importantly, the HLA-specific motifs remain broadly similar, but the stricter threshold retains only peptides with stronger predicted binding. The main effect of lowering the threshold is therefore to '''increase confidence in the retained HLA assignments at the cost of assigning many more peptides to Trash'''.
* more peptides enter the trash cluster; and
* the resulting HLA motifs are generally cleaner, but some real ligands may be discarded.


A '''higher %Rank threshold''' is less stringent:
At the 2% threshold, all five patients also have at least 10 Trash peptides, so MHCMotifDecon is able to generate a Trash motif for every patient.


* more peptides are assigned to HLA alleles;
A more stringent threshold may produce cleaner motifs, but it may also exclude genuine HLA ligands.
* fewer peptides enter the trash cluster; and
* weaker predicted binders and potentially contaminating peptides may be included.


Remember that for %Rank, '''smaller values indicate stronger predicted binding'''. The default MHCMotifDecon trash threshold for MHC class II is 20%.
*'''Q9: How many peptides are now associated with HLA-DRB4? Where did these peptides come from?'''


The exercise therefore illustrates the trade-off between '''sensitivity''' and '''specificity''': very stringent filtering produces cleaner assignments but risks losing true ligands, whereas permissive filtering retains more peptides but may introduce noise.
'''Answer:''' After including the secondary HLA-DR molecules, '''97 peptides''' were assigned to DRB3, DRB4 or DRB5.


*'''Q9: How many peptides are now associated with HLA-DRB4? Where did they come from?'''
[[Image:MHCMotifDecon_rank20_extendedHLA.png|center|800px]]


'''Answer:''' The exact DRB4 peptide count should be recorded from the second MHCMotifDecon run:
The assignments for each patient were:


DRB4_0101 = [number from your run]
RA1 DRB3_0202=19,DRB4_0101=8
RA2 DRB4_0101=14
RA3 DRB4_0101=12
RA4 DRB4_0101=1,DRB5_0101=7
RA5 DRB5_0101=36


The important observation is that adding '''DRB4_0101''' causes a subset of peptides to be reassigned to DRB4.
Overall:


'''Where did they come from?''' They generally come from peptides that, in the first analysis, had been:
DRB3 19 peptides
DRB4 35 peptides
DRB5 43 peptides


* assigned to one of the DRB1 molecules, particularly because DRB4 was not available as an alternative; or
These peptides were previously assigned either to one of the '''DRB1 alleles''' or to the '''Trash''' cluster because the corresponding secondary HLA-DR molecules were not included in the first analysis.
* placed in the trash cluster because neither DRB1*01:01 nor DRB1*04:01 provided a sufficiently good prediction.


'''Explanation:''' HLA-DRB1*04:01 is commonly linked with expression of '''HLA-DRB4*01:01'''. If DRB4 is genuinely expressed but is omitted from the HLA typing supplied to the algorithm, MHCMotifDecon has no opportunity to assign peptides to it.
Comparing the two runs shows:


This demonstrates an important principle of supervised motif deconvolution:
RA1 27 peptides reassigned: 12 from DRB1_0402, 7 from DRB1_1104 and 8 from Trash
RA2 14 peptides reassigned: 3 from DRB1_0101, 6 from DRB1_0401 and 5 from Trash
RA3 12 peptides reassigned: 1 from DRB1_0101, 10 from DRB1_0401 and 1 from Trash
RA4 8 peptides reassigned: 6 from DRB1_0401 and 2 from DRB1_1501
RA5 36 peptides reassigned: 28 from DRB1_0801, 7 from DRB1_1501 and 1 from Trash


'''The quality of the deconvolution depends on the completeness of the HLA typing supplied to the method.'''
For RA2, RA3 and RA5, only one new secondary allele was introduced, so the origin of the reassigned peptides can be determined directly from the change in counts.


Once DRB4 is included, peptides fitting its binding specificity can form a separate motif rather than being incorrectly associated with another allele or classified as trash.
For RA1, both '''DRB3_0202''' and '''DRB4_0101''' were added, and for RA4 both '''DRB4_0101''' and '''DRB5_0101''' were added. Therefore, the summary counts alone cannot determine exactly which previous DRB1 or Trash peptides were reassigned to each individual secondary allele.


DRB4 normally contributes a smaller peptide repertoire than the accompanying major DRB1 molecule, which also helps explain why its motif can be harder to identify with an unsupervised method.
The result demonstrates that including the secondary HLA-DR molecules substantially changes the deconvolution. In total, '''97 peptides''' that were previously attributed to DRB1 or Trash are instead predicted to be presented by DRB3, DRB4 or DRB5.


=== GibbsCluster ===
=== GibbsCluster ===
Line 135: Line 138:
*'''Q10: Does the number of clusters found by GibbsCluster correspond to the number of HLA alleles in your sample?'''
*'''Q10: Does the number of clusters found by GibbsCluster correspond to the number of HLA alleles in your sample?'''


'''Answer:''' Not necessarily. The number of sequence clusters identified by GibbsCluster does '''not have to equal the number of HLA molecules expressed by the sample'''.
'''Answer:''' No. When all five RA patients were combined, GibbsCluster tested solutions containing 1–5 clusters but selected a solution containing only '''one sequence motif'''.
 
This does not correspond to the number of HLA-DR molecules present in the samples. Across the five patients there are already '''six different DRB1 alleles''':


In this experiment, after taking DRB4 into account, there are three relevant HLA-DR specificities:
DRB1_0101
DRB1_0401
DRB1_0402
DRB1_0801
DRB1_1104
DRB1_1501


DRB1*01:01
If the secondary HLA-DR molecules identified in the previous exercise are also considered, the dataset additionally contains:
DRB1*04:01
DRB4*01:01


However, an unsupervised clustering analysis may recover only the strongest distinguishable motifs and may fail to resolve a smaller peptide population such as the DRB4 repertoire as an independent cluster.
DRB3_0202
DRB4_0101
DRB5_0101


'''Explanation:''' GibbsCluster does not know the patient's HLA genotype. It only examines patterns in the peptide sequences and asks how the sequences can best be partitioned into clusters.
Thus, several different HLA-DR specificities are present in the combined dataset, but GibbsCluster finds only '''one dominant motif'''.


Therefore:
This illustrates an important limitation of unsupervised motif deconvolution: the number of sequence clusters does not necessarily correspond directly to the number of HLA alleles. Different HLA molecules can have overlapping binding preferences, and smaller motifs may be hidden by larger, dominant peptide populations.


number of sequence clusters ≠ necessarily number of HLA alleles
In the selected one-cluster solution:


Two alleles with similar motifs may be merged into one cluster, a small allele-specific repertoire may not form a sufficiently strong independent cluster, or noisy peptides may affect the clustering solution.
Cluster 1 948 peptides
Outliers 94 peptides


This contrasts with MHCMotifDecon, where the known HLA molecules are explicitly supplied to the algorithm.


*'''Q11: Can you identify other differences between the GibbsCluster solution and the MHCMotifDecon solution?'''
*'''Q11: Can you identify other differences between the GibbsCluster solution and the MHCMotifDecon solution?'''


'''Answer:''' Yes. The most important difference is that '''MHCMotifDecon is supervised by HLA-binding predictions and known HLA typing, whereas GibbsCluster is an unsupervised sequence-clustering method'''.
'''Answer:''' Yes. The two methods produce quite different interpretations of the same immunopeptidomics data.
 
'''MHCMotifDecon''' uses the known HLA genotype of each patient together with peptide–HLA binding predictions. It therefore separates the peptides into motifs corresponding to specific HLA molecules, for example:
 
RA2 DRB1_0101,DRB1_0401,DRB4_0101
RA5 DRB1_0801,DRB1_1501,DRB5_0101
 
In contrast, '''GibbsCluster does not know which HLA alleles are present'''. It groups peptides only according to similarities in their amino-acid sequences. Consequently, when peptides from all five patients are combined, the different HLA-specific motifs are largely merged into a single dominant motif rather than being separated according to HLA allele.
 
GibbsCluster also identifies '''outliers''' rather than an allele-specific Trash group. In the selected solution, 94 peptides were classified as outliers.
 
Another important difference is that GibbsCluster first collapses the input to '''unique peptide sequences'''. In this analysis, 1072 unique sequences were read from the combined dataset. In contrast, MHCMotifDecon analysed the peptides separately according to their patient labels, allowing the same peptide to contribute independently to different patient-specific analyses.


'''Explanation:'''
Overall, MHCMotifDecon is better able to separate overlapping HLA motifs when the HLA genotype is known, whereas GibbsCluster attempts to discover motifs directly from the peptide sequences without prior HLA information.


With '''MHCMotifDecon''':
With '''MHCMotifDecon''':

Latest revision as of 22:40, 22 September 2026

Answers

Get the Data and Filter It

  • Q1: How many epitopes do you find with the initial filter?

Answer: 646 Epitopes.

  • Q2: Do you agree with calling these peptides epitopes? Why?

Answer: Not really. It would be more precise to call them MHC ligands or HLA-presented peptides.

Explanation: Mass spectrometry demonstrates that a peptide was isolated in association with an MHC molecule and therefore provides evidence that the peptide is naturally processed and presented. However, this does not by itself demonstrate that the peptide is recognized by a T-cell receptor or induces a T-cell response.

An epitope generally refers to a molecular structure that is recognized by the adaptive immune system. Therefore, an MHC ligand only becomes a demonstrated T-cell epitope when there is evidence of T-cell recognition.

  • Q3: How many eluted ligands did they find in this study?

Answer: Wang et al. found 2274 HLA-DR presented peptides

Be aware that the number of non-redundant presented peptides reported in the paper is not be identical to the IEDB number: 1,593. Why is that?

  • Q4: What kind of post-translational modifications are present?

+ DEAM(N2) and + CITR(R8) Deamination and Citrulation


  • Q5: Which alleles are expressed in the patient used for this assay?

Answer: Each RA sample has different alleles:

RA1    DRB1_0402,DRB1_1104
RA2    DRB1_0101,DRB1_0401
RA3    DRB1_0101,DRB1_0401
RA4    DRB1_0401,DRB1_1501
RA5    DRB1_0801,DRB1_1501

MHCMotifDecon

  • Q6: How many peptides were assigned to each of the alleles, and how many were assigned to the trash cluster?


RA1 DRB1_0402=141,DRB1_1104=71 Trash=18
RA2 DRB1_0101=193,DRB1_0401=245 Trash=52
RA3 DRB1_0101=124,DRB1_0401=301 Trash=11
RA4 DRB1_0401=51,DRB1_1501=16 Trash=3
RA5 DRB1_0801=101,DRB1_1501=31 Trash=7


A total of 1365 peptides were analysed. Although 1496 sequences were submitted, only peptides within the selected length range of 12–21 amino acids were included in the deconvolution.

  • Q7: Can you see any differences between the trash-cluster peptides and the peptides assigned to a specific allele? Explain.

Answer: Yes. Peptides assigned to specific HLA alleles show clear and reproducible binding motifs, with strong amino-acid preferences at particular positions of the peptide-binding core.

In contrast, the Trash clusters show much weaker and less defined sequence motifs. These peptides are predicted to bind poorly to all HLA alleles included for that patient and may therefore represent weak binders, contaminants, or peptides presented by an HLA molecule that was not included in the analysis.

For RA4 and RA5, only 3 and 7 peptides, respectively, were assigned to Trash. Because MHCMotifDecon requires at least 10 sequences to generate a logo, no Trash logo was produced for these patients.

  • Q8: What is the effect of changing the %Rank threshold?

Answer: Lowering the %Rank threshold from 20% to 2% makes the peptide assignment considerably more stringent. Peptides with a best predicted binding rank between 2% and 20% are no longer accepted as binders and are instead assigned to the Trash cluster.

The effect can be seen directly by comparing the two runs:

RA1 Trash: 18 → 51 (+33)
RA2 Trash: 52 → 91 (+39)
RA3 Trash: 11 → 34 (+23)
RA4 Trash: 3 → 11 (+8)
RA5 Trash: 7 → 37 (+30)

Across all five patients:

%Rank 20 91 peptides in Trash
%Rank 2 224 peptides in Trash

Thus, lowering the threshold to 2% moved 133 additional peptides into the Trash cluster.

Correspondingly, the number of peptides assigned to the HLA-DRB1 alleles decreased:

RA1 DRB1_0402: 141 → 120 DRB1_1104: 71 → 59
RA2 DRB1_0101: 193 → 177 DRB1_0401: 245 → 222
RA3 DRB1_0101: 124 → 118 DRB1_0401: 301 → 284
RA4 DRB1_0401: 51 → 44 DRB1_1501: 16 → 15
RA5 DRB1_0801: 101 → 77 DRB1_1501: 31 → 25

The total number of peptides assigned to an HLA allele therefore decreased from 1274 at the 20% threshold to 1141 at the 2% threshold.

Importantly, the HLA-specific motifs remain broadly similar, but the stricter threshold retains only peptides with stronger predicted binding. The main effect of lowering the threshold is therefore to increase confidence in the retained HLA assignments at the cost of assigning many more peptides to Trash.

At the 2% threshold, all five patients also have at least 10 Trash peptides, so MHCMotifDecon is able to generate a Trash motif for every patient.

A more stringent threshold may produce cleaner motifs, but it may also exclude genuine HLA ligands.

  • Q9: How many peptides are now associated with HLA-DRB4? Where did these peptides come from?

Answer: After including the secondary HLA-DR molecules, 97 peptides were assigned to DRB3, DRB4 or DRB5.

The assignments for each patient were:

RA1 DRB3_0202=19,DRB4_0101=8
RA2 DRB4_0101=14
RA3 DRB4_0101=12
RA4 DRB4_0101=1,DRB5_0101=7
RA5 DRB5_0101=36

Overall:

DRB3 19 peptides
DRB4 35 peptides
DRB5 43 peptides

These peptides were previously assigned either to one of the DRB1 alleles or to the Trash cluster because the corresponding secondary HLA-DR molecules were not included in the first analysis.

Comparing the two runs shows:

RA1 27 peptides reassigned: 12 from DRB1_0402, 7 from DRB1_1104 and 8 from Trash
RA2 14 peptides reassigned: 3 from DRB1_0101, 6 from DRB1_0401 and 5 from Trash
RA3 12 peptides reassigned: 1 from DRB1_0101, 10 from DRB1_0401 and 1 from Trash
RA4 8 peptides reassigned: 6 from DRB1_0401 and 2 from DRB1_1501
RA5 36 peptides reassigned: 28 from DRB1_0801, 7 from DRB1_1501 and 1 from Trash

For RA2, RA3 and RA5, only one new secondary allele was introduced, so the origin of the reassigned peptides can be determined directly from the change in counts.

For RA1, both DRB3_0202 and DRB4_0101 were added, and for RA4 both DRB4_0101 and DRB5_0101 were added. Therefore, the summary counts alone cannot determine exactly which previous DRB1 or Trash peptides were reassigned to each individual secondary allele.

The result demonstrates that including the secondary HLA-DR molecules substantially changes the deconvolution. In total, 97 peptides that were previously attributed to DRB1 or Trash are instead predicted to be presented by DRB3, DRB4 or DRB5.

GibbsCluster

  • Q10: Does the number of clusters found by GibbsCluster correspond to the number of HLA alleles in your sample?

Answer: No. When all five RA patients were combined, GibbsCluster tested solutions containing 1–5 clusters but selected a solution containing only one sequence motif.

This does not correspond to the number of HLA-DR molecules present in the samples. Across the five patients there are already six different DRB1 alleles:

DRB1_0101
DRB1_0401
DRB1_0402
DRB1_0801
DRB1_1104
DRB1_1501

If the secondary HLA-DR molecules identified in the previous exercise are also considered, the dataset additionally contains:

DRB3_0202
DRB4_0101
DRB5_0101

Thus, several different HLA-DR specificities are present in the combined dataset, but GibbsCluster finds only one dominant motif.

This illustrates an important limitation of unsupervised motif deconvolution: the number of sequence clusters does not necessarily correspond directly to the number of HLA alleles. Different HLA molecules can have overlapping binding preferences, and smaller motifs may be hidden by larger, dominant peptide populations.

In the selected one-cluster solution:

Cluster 1	948 peptides
Outliers	94 peptides


  • Q11: Can you identify other differences between the GibbsCluster solution and the MHCMotifDecon solution?

Answer: Yes. The two methods produce quite different interpretations of the same immunopeptidomics data.

MHCMotifDecon uses the known HLA genotype of each patient together with peptide–HLA binding predictions. It therefore separates the peptides into motifs corresponding to specific HLA molecules, for example:

RA2	DRB1_0101,DRB1_0401,DRB4_0101
RA5	DRB1_0801,DRB1_1501,DRB5_0101

In contrast, GibbsCluster does not know which HLA alleles are present. It groups peptides only according to similarities in their amino-acid sequences. Consequently, when peptides from all five patients are combined, the different HLA-specific motifs are largely merged into a single dominant motif rather than being separated according to HLA allele.

GibbsCluster also identifies outliers rather than an allele-specific Trash group. In the selected solution, 94 peptides were classified as outliers.

Another important difference is that GibbsCluster first collapses the input to unique peptide sequences. In this analysis, 1072 unique sequences were read from the combined dataset. In contrast, MHCMotifDecon analysed the peptides separately according to their patient labels, allowing the same peptide to contribute independently to different patient-specific analyses.

Overall, MHCMotifDecon is better able to separate overlapping HLA motifs when the HLA genotype is known, whereas GibbsCluster attempts to discover motifs directly from the peptide sequences without prior HLA information.

With MHCMotifDecon:

  • the HLA alleles expressed by the sample are supplied beforehand;
  • each peptide is evaluated using allele-specific binding predictions;
  • clusters are directly labelled with specific HLA molecules;
  • peptides that do not bind any supplied allele sufficiently well can be placed in a trash cluster; and
  • relatively small allele-specific peptide populations may still be detected because the algorithm already knows which alleles to test.

With GibbsCluster:

  • no HLA genotype information is required;
  • peptides are grouped according to similarities in their sequence motifs;
  • the resulting clusters are not intrinsically labelled as particular HLA alleles;
  • allele identities have to be inferred afterwards by comparing the motifs with known HLA-binding motifs; and
  • weak or small motifs can be merged with larger clusters or may not emerge as independent clusters.

Therefore, the two approaches answer slightly different questions.

MHCMotifDecon asks: Given the HLA molecules that I know are present, which HLA molecule most likely presented each peptide?

GibbsCluster asks: Without knowing which HLA molecules are present, how many different sequence patterns can I detect in this peptide dataset?

Using both approaches is useful because agreement between them gives additional confidence in the inferred motifs, while disagreement can reveal incomplete HLA typing, low-abundance HLA molecules, overlapping binding specificities, or contaminants.