<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?action=history&amp;feed=atom&amp;title=Regular_expressions</id>
	<title>Regular expressions - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?action=history&amp;feed=atom&amp;title=Regular_expressions"/>
	<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;action=history"/>
	<updated>2026-10-04T11:33:04Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=348&amp;oldid=prev</id>
		<title>WikiSysop: /* Required course material for the lesson */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=348&amp;oldid=prev"/>
		<updated>2026-09-03T08:29:58Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Required course material for the lesson&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 10:29, 3 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l12&quot;&gt;Line 12:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 12:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;PDF: [https://teaching.healthtech.dtu.dk/material/22116/regular-expressions-cheat-sheet-v2.pdf Regular Expressions Cheat Sheet]&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;PDF: [https://teaching.healthtech.dtu.dk/material/22116/regular-expressions-cheat-sheet-v2.pdf Regular Expressions Cheat Sheet]&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;WWW: [http://regex101.com/ Web page where you can test your regular expressions]&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;WWW: [http://regex101.com/ Web page where you can test your regular expressions]&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Resource: [[Example code - exam form]]&amp;lt;br&amp;gt;&lt;/ins&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Subjects covered ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Subjects covered ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=347&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=347&amp;oldid=prev"/>
		<updated>2026-09-02T12:58:57Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:58, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l25&quot;&gt;Line 25:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 25:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Exercise 6 to 8 has strong taste of something I would do at an exam. It is also an interesting beginning of the making a HIV vaccine. The data is real and the methods are real.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Exercise 6 to 8 has strong taste of something I would do at an exam. It is also an interesting beginning of the making a HIV vaccine. The data is real and the methods are real.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Make a function that accepts a string (X) as parameter. Use regular expressions (RE) to determine if the string is a number and outputs either &quot;X is a number&quot; or &quot;X is not a number&quot; depending on the string content. The goal is to do this with a SINGLE regex.&amp;lt;br&amp;gt; These should all be considered as numbers: &quot;4&quot;   &quot;-7&quot;   &quot;0.656&quot;   &quot;-67.35555&quot;&amp;lt;br&amp;gt; These are not numbers: &quot;5.&quot;  &quot;56F&quot;  &quot;.32&quot;  &quot;-.04&quot;  &quot;1+1&quot;&amp;lt;br&amp;gt; Note: The program is very simple, but it is likely the most difficult regular expression, you will have to make in this set of exercises. Perhaps you should do the following exercises before attempting this one - just to get some experience first.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Make a function that accepts a string (X) as parameter. Use regular expressions (RE) to determine if the string is a number and outputs either &quot;X is a number&quot; or &quot;X is not a number&quot; depending on the string content. The goal is to do this with a SINGLE regex.&amp;lt;br&amp;gt; These should all be considered as numbers: &quot;4&quot;   &quot;-7&quot;   &quot;0.656&quot;   &quot;-67.35555&quot;&amp;lt;br&amp;gt; These are not numbers: &quot;5.&quot;  &quot;56F&quot;  &quot;.32&quot;  &quot;-.04&quot;  &quot;1+&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;1&quot; &quot;1-&lt;/ins&gt;1&quot;&amp;lt;br&amp;gt; Note: The program is very simple, but it is likely the most difficult regular expression, you will have to make in this set of exercises. Perhaps you should do the following exercises before attempting this one - just to get some experience first.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Regex is often used for validation. This time make a function that takes a string as input, and checks if it is in a proper format for an email address. Output is &amp;quot;X is a valid email address&amp;quot; or &amp;quot;X is not a valid email address&amp;quot;. Quite similar to the previous exercise, just a different thing to validate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Regex is often used for validation. This time make a function that takes a string as input, and checks if it is in a proper format for an email address. Output is &amp;quot;X is a valid email address&amp;quot; or &amp;quot;X is not a valid email address&amp;quot;. Quite similar to the previous exercise, just a different thing to validate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; for the reading part. The verification here means that the function prints &amp;quot;filename is DNA fasta&amp;quot; or &amp;quot;filename is protein fasta&amp;quot; if the file is successfully verified for either dna or protein sequence, and &amp;quot;filename is not fasta&amp;quot; if unsuccessfully verified. Test the function with &amp;#039;&amp;#039;dna7.fsa&amp;#039;&amp;#039; and &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up.&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; for the reading part. The verification here means that the function prints &amp;quot;filename is DNA fasta&amp;quot; or &amp;quot;filename is protein fasta&amp;quot; if the file is successfully verified for either dna or protein sequence, and &amp;quot;filename is not fasta&amp;quot; if unsuccessfully verified. Test the function with &amp;#039;&amp;#039;dna7.fsa&amp;#039;&amp;#039; and &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up.&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=346&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=346&amp;oldid=prev"/>
		<updated>2026-09-02T12:43:25Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:43, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l30&quot;&gt;Line 30:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 30:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you return. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you return. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. Make a function that accepts an input &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;fasta &lt;/del&gt;filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. Make a function that accepts an input &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;SwissProt &lt;/ins&gt;filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other. Create a function that accepts an input epitope file and an output epitope file to perform the job, possibly also creating other helper functions in the process.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance, [[Functions]] exercise 1. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other. Create a function that accepts an input epitope file and an output epitope file to perform the job, possibly also creating other helper functions in the process.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance, [[Functions]] exercise 1. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=345&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=345&amp;oldid=prev"/>
		<updated>2026-09-02T12:40:38Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:40, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l29&quot;&gt;Line 29:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 29:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; for the reading part. The verification here means that the function prints &amp;quot;filename is DNA fasta&amp;quot; or &amp;quot;filename is protein fasta&amp;quot; if the file is successfully verified for either dna or protein sequence, and &amp;quot;filename is not fasta&amp;quot; if unsuccessfully verified. Test the function with &amp;#039;&amp;#039;dna7.fsa&amp;#039;&amp;#039; and &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up.&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; for the reading part. The verification here means that the function prints &amp;quot;filename is DNA fasta&amp;quot; or &amp;quot;filename is protein fasta&amp;quot; if the file is successfully verified for either dna or protein sequence, and &amp;quot;filename is not fasta&amp;quot; if unsuccessfully verified. Test the function with &amp;#039;&amp;#039;dna7.fsa&amp;#039;&amp;#039; and &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up.&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastaread()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &#039;&#039;alignment.fsa&#039;&#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;print&lt;/del&gt;. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &#039;&#039;alignment.fsa&#039;&#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;return&lt;/ins&gt;. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=344&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=344&amp;oldid=prev"/>
		<updated>2026-09-02T12:11:09Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:11, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l32&quot;&gt;Line 32:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 32:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance, [[Functions]] exercise 1. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;. Create a function that accepts an input epitope file and an output epitope file to perform the job, possibly also creating other helper functions in the process&lt;/ins&gt;.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance, [[Functions]] exercise 1. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=343&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=343&amp;oldid=prev"/>
		<updated>2026-09-02T12:05:54Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:05, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l32&quot;&gt;Line 32:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 32:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV be creating a function that accepts an input fasta filename and and output epitope file name. The functions reads the input &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and creates a single set consisting of all possible epitopes for all sequences. Save the epitopes in the output file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitopes left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;, [[Functions]] exercise 1&lt;/ins&gt;. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=342&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=342&amp;oldid=prev"/>
		<updated>2026-09-02T12:02:45Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:02, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l31&quot;&gt;Line 31:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 31:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. You did something similar back in the [[Stateful parsing]] lesson, exercise 6.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Read &lt;/del&gt;the &#039;&#039;HIVenv.fsa&#039;&#039; fasta file and &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;create &lt;/del&gt;a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &#039;&#039;HIVepitopes.txt&#039;&#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;be creating a function that accepts an input fasta filename and and output epitope file name&lt;/ins&gt;. &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;The functions reads &lt;/ins&gt;the &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;input &lt;/ins&gt;&#039;&#039;HIVenv.fsa&#039;&#039; fasta file and &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;creates &lt;/ins&gt;a single set consisting of all possible epitopes for all sequences. Save the epitopes in the &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;output &lt;/ins&gt;file &#039;&#039;HIVepitopes.txt&#039;&#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;epitope &lt;/del&gt;left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &#039;&#039;HIVepitopes.txt&#039;&#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;epitopes &lt;/ins&gt;left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &quot;surviving&quot; epitopes in the file &#039;&#039;HIVepitopesML.txt&#039;&#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=341&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=341&amp;oldid=prev"/>
		<updated>2026-09-02T11:57:39Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 13:57, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l30&quot;&gt;Line 30:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 30:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;YouThis job is not new to you. and you can use your &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function if you want to&lt;/del&gt;.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. Make a function that accepts an input fasta filename and and output fasta filename, and uses regular repressions to extract the ID and the protein sequence from each entry in the input file and save them in an output fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;You did something similar back in the [[Stateful parsing]] lesson, exercise 6&lt;/ins&gt;.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitope left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitope left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=340&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=340&amp;oldid=prev"/>
		<updated>2026-09-02T11:54:18Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 13:54, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l30&quot;&gt;Line 30:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 30:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &amp;#039;&amp;#039;dnanoise.fsa&amp;#039;&amp;#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;You are welcome to use the &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function in other exercises if appropriate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &amp;#039;&amp;#039;alignment.fsa&amp;#039;&amp;#039; is a protein alignment of part of the insulin gene from different organisms. Make a function that reads the input fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Using &lt;/del&gt;regular repressions &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;you must &lt;/del&gt;extract the ID and the protein sequence from each entry and save them in &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;a &lt;/del&gt;fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;This &lt;/del&gt;job is not new to you and you can use your &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function if you want to.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &#039;&#039;HIVenvelope.txt&#039;&#039;. &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Make a function that accepts an input fasta filename and and output fasta filename, and uses &lt;/ins&gt;regular repressions &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;to &lt;/ins&gt;extract the ID and the protein sequence from each entry &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;in the input file &lt;/ins&gt;and save them in &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;an output &lt;/ins&gt;fasta file named &#039;&#039;HIVenv.fsa&#039;&#039;. &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;YouThis &lt;/ins&gt;job is not new to you&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;. &lt;/ins&gt;and you can use your &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function if you want to.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitope left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# You must prepare the epitopes in &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; for machine learning. A common ML problem is when you have many data points (here epitopes) that look like each other. This introduces an unwanted bias in the ML predictions. Your job is to eliminate epitopes that are too similar using the Hobohm-1 algorithm. Hobohm-1 works like this: Look a the first epitope in the list. Compare it sequentially with the rest of the epitopes. If an epitope is too similar with the first, throw it away. Now no epitopes looks like the first. Proceed to the second epitope and repeat the comparing and possibly throwing away of the subsequent epitopes. Proceed to the third epitope in the list. Repeat this pattern until you have reached the end of the list. Now all epitope left are dissimilar to each other.&amp;lt;br&amp;gt;How to determine if two epitopes are too similar? Easy - you earlier learned about the Hamming distance. Just compute the Hamming distance between two epitopes and if the distance is 3 or less, they are too similar. Save the &amp;quot;surviving&amp;quot; epitopes in the file &amp;#039;&amp;#039;HIVepitopesML.txt&amp;#039;&amp;#039;.&amp;lt;br&amp;gt;Hint: Think about the Hobohm-1 algorithm before you implement it. You can easily run into problems.&amp;lt;br&amp;gt;You should see a reduction from 15909 epitopes to 3842 in 30-60 seconds.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Exercises for extra practice ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
	<entry>
		<id>https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=339&amp;oldid=prev</id>
		<title>WikiSysop: /* Exercises to be handed in */</title>
		<link rel="alternate" type="text/html" href="https://teaching.healthtech.dtu.dk:443/22116/index.php?title=Regular_expressions&amp;diff=339&amp;oldid=prev"/>
		<updated>2026-09-02T11:49:08Z</updated>

		<summary type="html">&lt;p&gt;&lt;span dir=&quot;auto&quot;&gt;&lt;span class=&quot;autocomment&quot;&gt;Exercises to be handed in&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 13:49, 2 September 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l27&quot;&gt;Line 27:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 27:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Make a function that accepts a string (X) as parameter. Use regular expressions (RE) to determine if the string is a number and outputs either &amp;quot;X is a number&amp;quot; or &amp;quot;X is not a number&amp;quot; depending on the string content. The goal is to do this with a SINGLE regex.&amp;lt;br&amp;gt; These should all be considered as numbers: &amp;quot;4&amp;quot;   &amp;quot;-7&amp;quot;   &amp;quot;0.656&amp;quot;   &amp;quot;-67.35555&amp;quot;&amp;lt;br&amp;gt; These are not numbers: &amp;quot;5.&amp;quot;  &amp;quot;56F&amp;quot;  &amp;quot;.32&amp;quot;  &amp;quot;-.04&amp;quot;  &amp;quot;1+1&amp;quot;&amp;lt;br&amp;gt; Note: The program is very simple, but it is likely the most difficult regular expression, you will have to make in this set of exercises. Perhaps you should do the following exercises before attempting this one - just to get some experience first.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Make a function that accepts a string (X) as parameter. Use regular expressions (RE) to determine if the string is a number and outputs either &amp;quot;X is a number&amp;quot; or &amp;quot;X is not a number&amp;quot; depending on the string content. The goal is to do this with a SINGLE regex.&amp;lt;br&amp;gt; These should all be considered as numbers: &amp;quot;4&amp;quot;   &amp;quot;-7&amp;quot;   &amp;quot;0.656&amp;quot;   &amp;quot;-67.35555&amp;quot;&amp;lt;br&amp;gt; These are not numbers: &amp;quot;5.&amp;quot;  &amp;quot;56F&amp;quot;  &amp;quot;.32&amp;quot;  &amp;quot;-.04&amp;quot;  &amp;quot;1+1&amp;quot;&amp;lt;br&amp;gt; Note: The program is very simple, but it is likely the most difficult regular expression, you will have to make in this set of exercises. Perhaps you should do the following exercises before attempting this one - just to get some experience first.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Regex is often used for validation. This time make a function that takes a string as input, and checks if it is in a proper format for an email address. Output is &amp;quot;X is a valid email address&amp;quot; or &amp;quot;X is not a valid email address&amp;quot;. Quite similar to the previous exercise, just a different thing to validate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Regex is often used for validation. This time make a function that takes a string as input, and checks if it is in a proper format for an email address. Output is &amp;quot;X is a valid email address&amp;quot; or &amp;quot;X is not a valid email address&amp;quot;. Quite similar to the previous exercise, just a different thing to validate.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &#039;&#039;&#039;fastaread()&#039;&#039;&#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &#039;&#039;&#039;fastaread()&#039;&#039;&#039; for the reading part. The verification here means that the function prints &quot;filename is DNA fasta&quot; or &quot;filename is protein fasta&quot; if the file is successfully verified for either dna or protein sequence, and &quot;filename is not fasta&quot; if unsuccessfully verified. Test the function with &#039;&#039;dna7.fsa&#039;&#039; and &#039;&#039;dnanoise.fsa&#039;&#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &#039;&#039;&#039;fastaread()&#039;&#039;&#039; function in exercise 5. Start with inserting that function in the python file at the designated place. Now a function that accepts a filename as parameter and reads and verifies a fasta file. Use your  &#039;&#039;&#039;fastaread()&#039;&#039;&#039; for the reading part. The verification here means that the function prints &quot;filename is DNA fasta&quot; or &quot;filename is protein fasta&quot; if the file is successfully verified for either dna or protein sequence, and &quot;filename is not fasta&quot; if unsuccessfully verified. Test the function with &#039;&#039;dna7.fsa&#039;&#039; and &#039;&#039;dnanoise.fsa&#039;&#039;. You can find a description of fasta format in [[Biological knowledge needed in the course]]. You are expected to know which symbols are used for DNA and protein sequence - otherwise look it up&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;.&amp;lt;br&amp;gt;You are welcome to use the &#039;&#039;&#039;fastaread()&#039;&#039;&#039; function in other exercises if appropriate&lt;/ins&gt;.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &#039;&#039;dnanoise.fsa&#039;&#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989).&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# In the [[Functions]] lesson you made a &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function in exercise 6. Start with inserting that function in the python file at the designated place. Building on your experience with the previous exercise, make a function that takes as input an input-fasta-filename and and output-fastafilename. The function reads the input fasta file, discard entries that do not conform to DNA or protein sequence, and rewrite (using your &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function) the acceptable entries in the output fasta file, in such a way that the normal 60 chars per line is followed with no spaces in between. The function must inform the user how many entries was kept and how many discarded. Hint: Test on &#039;&#039;dnanoise.fsa&#039;&#039;, which contains 3 entries that should be discarded (V00179, J00265, J02989)&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;.&amp;lt;br&amp;gt;You are welcome to use the &#039;&#039;&#039;fastawrite()&#039;&#039;&#039; function in other exercises if appropriate&lt;/ins&gt;.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &#039;&#039;alignment.fsa&#039;&#039; is a protein alignment of part of the insulin gene from different organisms. &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Read &lt;/del&gt;the fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# [[File:consensus.png|50px|right]] In the file &#039;&#039;alignment.fsa&#039;&#039; is a protein alignment of part of the insulin gene from different organisms. &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Make a function that reads &lt;/ins&gt;the &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;input &lt;/ins&gt;fasta file and determine the consensus sequence of the alignment, which you print. The consensus sequence is simply the most frequent occurring amino acid on each position.&amp;lt;br&amp;gt;Consensus: MALWMRLLPLLALLALWEPDPAGAFVNGHLCGSHLVEALYLVCGERGFFYTPKSRREVEDPGVGGLELGGGP&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Using regular repressions you must extract the ID and the protein sequence from each entry and save them in a fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. This job is not new to you and you can use your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function if you want to.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# All HIV envelope proteins from various HIV strains in SwissProt have been identified and collected in the file &amp;#039;&amp;#039;HIVenvelope.txt&amp;#039;&amp;#039;. Using regular repressions you must extract the ID and the protein sequence from each entry and save them in a fasta file named &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039;. This job is not new to you and you can use your &amp;#039;&amp;#039;&amp;#039;fastawrite()&amp;#039;&amp;#039;&amp;#039; function if you want to.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;# Continuing the investigation in HIV. Read the &amp;#039;&amp;#039;HIVenv.fsa&amp;#039;&amp;#039; fasta file and create a single set consisting of all possible epitopes for all sequences. Save the epitopes in the file &amp;#039;&amp;#039;HIVepitopes.txt&amp;#039;&amp;#039; - one epitope per line. An epitope is simply a k-mer 9 residues long, which can possibly elicit a immune system response. So save all unique 9-mers in the sequences in the file. To generate a file that you can verify against my file, you must sort the epitopes alphabetically before saving them.&amp;lt;br&amp;gt;&amp;lt;br&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>WikiSysop</name></author>
	</entry>
</feed>