Disease Motifs
An Excursion into the Field of Bioinformatics

Disease Motifs

Introduction

This site is dedicated to some bioinformatics work which looks at the proteins associated with a particular disease such as Parkinson’s Disease, Alzheimer’s Disease, certain types of cancer, and others etc. It looks at the proteins that have been seen to have an important role in the disease and looks for small amino acid sequences that might be significant – disease motifs. There are different ways to analyse the proteins, such as splitting the group up into sets and also considering from which chromosome these proteins are coded.

The research work that can be seen here is done in my spare time and I like to think that at some point I might find something relevant to a particular disease; however, this has not proven to be an easy task.

For a small introduction to the work on this site please watch the following video 001 Introduction on Patreon. This video is free to watch.


Primary Disease Dataset

Parkinson's Disease Protein Dataset

The Parkinson's Disease Protein Dataset has been the primary dataset I have been involved with for a number of years; it was one of the first datasets that I studied whilst developing this bioinformatics technique. I might change this focus shortly for one reason or another, especially if I managed to get some sponsorship, in which case the material for such work would appear here. Other protein disease datasets are listed below.

To see all of the protein sequence analysis on the proteins that are associated with Parkinson's disease (the Parkinson's Disease Protein Dataset) click on the following link: Parkinson's disease protein dataset analysis. For more information see the video 024 Disease Motifs Analysis of the Parkinson's Disease Protein Dataset on Patreon. This video is free to watch.

When you have clicked on the above link you will then need to click on the "protein sequence analysis" link for various proteins on the subsequent pages, this will give you some idea of how the data outputs will look that are produced by the Perl scripts.





Other Disease Protein Datasets

Adenocarcinoma Protein Dataset

The data from the Adenocarcinoma Protein Dataset is found here at Adenocarcinoma Protein Dataset.
For more information see the video 039 Adenocarcinoma 2026 on Patreon. This video is free to watch.





Alzheimer's Disease Protein Dataset

The data from the Alzheimer's Disease Protein Dataset is found here at Alzheimer's Disease Protein Dataset. For more information see the video 014 Alzheimer's Disease Protein Dataset 2025 on Patreon. This video is free to watch.





Amyloidosis Protein Dataset

A small set of proteins involved in Amyloidosis, click on the following link to see the data analysis: Amyloidosis Protein Dataset. For more information see the video 015 Amyloidosis Protein Dataset 2025 on Patreon. This video is free to watch.




Disease Motifs - An Excursion into the Field of Bioinformatics

Disease Motifs
An Excursion into the Field of Bioinformatics

PAPERBACK

Disease Motifs
An Excursion into the Field of Bioinformatics

eBook

The whole basis of this work is to use proteomic analysis to investigate the causes of some diseases. As complicated as this sounds, this just means looking at the amino acid sequences of proteins that might have some association with a particular disease: we could just say that this is protein sequence analysis, trying to find some significance in the amino acid sequence by comparing with other amino acid sequences from other proteins. If you are interested in the ideas behind all of this then you will need to buy the book.


Support My Work

A picture of me taken in 2005 (many years ago) and yes I still wear that shirt

If you are interested in sponsoring me to work on a particular topic, or just interested in sponsoring this type of work in general without specifying any particular pathology, please let me know - but this won't be cheap.

While it might be very desirable, as a sponsor, to donate money to a project, or, as a researcher, to receive funds to conduct research into a particular topic for years in advance as it might take years to find anything of value, I think that that it would be better if it is funded for only one year at a time. Firstly, this would give me the chance to decide if it was worthwhile continuing further research regarding a particular topic, or not. Secondly, it will give the sponsor time to decide if they thought the work is of any value and whether they thought that they were getting good value for their money. It might be that after one year I was no longer interested or capable (for whatever reason) of continuing the work, or it might be that the sponsor was no longer interested in my work and no longer saw any reason for supporting it. We would both then have the right to terminate our agreement. In fact, the project would automatically end after a one-year period. The fee for one year’s work would have to be payable in advance as I would need to makes changes to my life's routines to accommodate the research.

You should know that I would not intend to work myself to death in this work period, working on average around twenty-one hours per week with a six-week annual leave entitlement. And if I die prematurely then that would be that. And once money has been paid there would be no chance of a refund.

I will only undertake to do one such project at any one time, as I have other interests, activities and obligations other than bioinformatics. This is a Wild Card donation as there is no guarantee that this work would discover anything of medical significance even after ten years, but then again it might - the choice is yours as to whether you think it is worth a gamble or not. If you are interested send me an email.

If you are interested in finding out more about this work and scripts used, you could join me on DiseaseMotifs711 Patreon site which will give you access to specific screen-capture-videos and Perl scripts which you may find of interest. See links below.


Perl Scripts, Video Tutorials, and Explanations

Here is an index of screen-capture explanatory videos that can be viewed on Patreon - there is membership fee to pay for most of these but there are some that are for free (I will increase the free-viewing ones over time). Click on the links (highlighted in red) to go to the Patreon site for the selected video.

If any of the links are not working please let me know, so that I can fix these.

There are also all of the Perl scripts that can be downloaded - this will save you from having to write them out.

This is not a complete list as I may not necessarily update this list if new material is released - YOU WILL NEED TO VISIT THE PATREON SITE at DiseaseMotifs711 Patreon site .


001 A Short Introduction
This is a short introduction to my work and to my website Disease Motifs. It is my first attempt at publishing a short podcast using a screen recorder. This worked okay but there are two occasions when the screen freezes for a few milliseconds, but I decided to release this as the rest is fine. There are a couple of other errors in that when I am talking about the ranking of the Parkin protein I have this ranked at 140 (140 times the word "parkinson" was counted in the protein datasheet) and in fact when I do a "Find" search in the datasheet you can see 141; I'm not sure why this is but I may decide to look at this at some future date if I have the time. Please watch the following video 001 Introduction on Patreon. This video is free to watch.

Another error was that I had mentioned that I would explain why cancer cropped up in the Parkinson's research, but I forgot to give an example in this video, this is because it had taken me several attempts to get this video right and I had mentions this in one of the previous deleted precursors of this video and forgot to mention in this video. It doesn't matter really as I will do another video at some point regarding this.


002 Information about some of the Polymorphisms
This is the second video and offers an explanation about some of the polymorphisms. In it I explain that not all polymorphisms are picked out of the datasheets only those that seem to be associated with a pathology in some way. Also I highlight how on some occasions the Perl script that I wrote misses the key phrase that explain an association with a pathology - usually selecting the letter "a" as in "in a colorectal cancer sample".
Out of the 129 proteins in the Parkinson's Disease dataset 25 have no "accepted" polymorphisms, leaving just 104 that do have polymorphisms, and most of these are not polymorphisms associated with Parkinson's Disease.


003 Set-Theory
An explanation about set-theory used in Disease Motifs. A short twenty minute video explaning the use of Set Theory in some previously published work This also explains the reasons behind the use of genetic chromosomal and loci data on the set-theory page of the site.


004 Predicted Onco-Motif
This is about a predicted onco-motif found in the Platelet-derived growth factor receptor beta (PGFRB_HUMAN) protein. Several years ago there were two maxima of interest found at sequences 169, and 178.  Sequence 177 was also later seen as importantly connected.

Sequence 178 (VWSFGILLWE) returned 250 proteins of many different species but there were only eight viral proteins present, all from viruses that have been known to cause cancer albeit in non-human hosts e.g. mice, chickens.  Forty human proteins were returned; the majority having some association with growth. This motif may extend to seq 177 (TTLSDVWSFG).


005 Park
Explanation of the Park set, which supposedly consists of proteins with polymorphisms associated with Parkinson's Disease, where there is the "erroneously caught" protein (PACRG_HUMAN) that has no polymorphisms.
Erratum:
I did say Pestis vaccine in the video when I meant to say Pestis protein.


006 Downloading the Swiss-Prot Protein Databases from Uniprot in 2025
How to download the Swiss-Prot protein databases from the Uniprot website as of Wednesday 12th of February 2025. You will need these databases to duplicate the work seen on the Disease Motifs website but you might also want to do some of your own work using the datasets.
Erratum:
I did call the "date" text file a folder near the end of the video.

You might find this link to the User Manual on the expasy site https://web.expasy.org/docs/userman.html useful as it explains the annotation of the Swiss-Prot datasheets. Useful search terms to find this via a search engine are "Expasy", "Swiss-Prot", and "User Manual".

To download the Swiss-Prot databases that are used in this work go to https://www.uniprot.org/help/downloads

To download the gentic information from the HGNC HGNC data


007 Downloading the blastall and formatdb executable files
Downloading the blastall.exe and formatdb.exe from the National Center for Biotechnology Information section on the National Library of Medicine website (an American website).

If you want the blastall and formatdb executable files which are used with the Perl scripts then try this link BLAST and FORMAT programmes which are used with the Perl scripts. There will be some warnings about this not being safe but if you want these you will have to download them. These are legacy 32 bit programmes in the NOTSUPPORTED section and are version 2.2.19. These and can be downloaded from the National Center for Biotechnology Information NCBI section on the National Institute of Health NIH (American) website. Don't OPEN the file when they have been downloaded, but just double-click on them - for more information and details please refer to the books on Amazon (see below for links), and the videos on Patreon (see below for links).


008 Installing the Strawberry Perl Interpreter
Downloading and installing the Strawberry Perl Interpreter. This is needed to run Perl scripts.


009 Using the Command Prompt to run Perl Scripts
Using the Command Prompt to run Perl scripts.


010 The Sampler Scripts
This compressed folder contains two Perl scripts that "sample" the seven main databases: all, human, bacteria, fungi, viruses, archaea, and plants. The first one is the "samplerFirst100.pl" script, which selects the first one hundred entries in the database; and the second one is the "sampler1in1000.pl" script, which samples one entry in every 1000 entries in the database.

Link to the Sampler Perl Scripts
This compressed folder contains two Perl scripts that "sample" the seven main databases: all, human, bacteria, fungi, viruses, archaea, and plants. The first one is the "samplerFirst100.pl" script, which selects the first one hundred entries in the database; and the second one is the "sampler1in1000.pl" script, which samples one entry in every 1000 entries in the database.

You will need to extract these from the compressed folder to be able to use them. Obviously you will need to have installed the Perl Interpreter for these to work. Please see the relevant video about installing the interpreter and also the Sampler video.


011 OHfilegrab
OHfilegrab grabs every datasheet entry with a specified species named in the Organism Host line on a Swiss-Prot protein datasheet. The optional OH line is specific for viruses and states any hosts that the virus will infect. I used the OHfilegrab to make the humanVirus (note the capital "V") dat file.
Link to the OHfilegrab Perl script on Patreon


012 OSfilegrab
OSfilegrab pulls out any datasheet with a specified species name on the OS (Organism Species) line. The OS line specifies the organism that the protein from the datasheet has come from. I used this programme to create a rickettsia dat file.
Link to the OSfilegrab Perl script on Patreon


013 Parkinson's Disease Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Parkinson's Disease Protein Dataset from the Swiss-Prot human protein database. There are 136 proteins in the Parkinson's Disease Protein Dataset 2025. Two proteins were deleted from the original 138 proteins as these were deemed to be false positives.
The DiseaseGrab scripts are:
master.pl
masterCANCER.pl (not to be used unless you have already created a "cancer" dat file)
DiseaseGrab.pl
fastaD.pl
fastaDeliver.pl
FTvariant.pl
This can be downloaded in the form of a zip file and needs to be extracted and placed "correctly" in a directory before being used (see video).
Link to the DiseaseGrab Perl scripts on Patreon


014 Alzheimer's Disease Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Alzheimer's Disease Protein Dataset from the Swiss-Prot human protein database. There are 114 protiens in the Alzheimer's Disease Protein Dataset. For more information see the video 014 Alzheimer's Disease Protein Dataset 2025 on Patreon. This video is free to watch.


015 Amyloidosis Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Amyloidosis Protein Dataset from the Swiss-Prot human protein database. There are 21 proteins in the Amyloidosis Protein Dataset. For more information see the video 015 Amyloidosis Protein Dataset 2025 on Patreon. This video is free to watch.


016 Cancer Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Cancer Protein Dataset from the Swiss-Prot human protein database. There are 2,129 proteins in the Cancer Protein Dataset.


017 Viral Cancer Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Viral Cancer Protein Dataset from the Swiss-Prot viruses protein database. There are 131 proteins in the Viral Cancer Protein Dataset.


018 Cancer Derivative Protein Datasets 2025
Using the DiseaseGrab scripts to extract the Cancer Derivative Protein Datasets from the cancer protein dat file.


019 Leukemia Protein Dataset 2025
Using the DiseaseGrab scripts to extract Leukemia Protein Dataset from the Swiss-Prot human protein database. There are 460 proteins in the Leukemia Protein Dataset.


020 Viral Leukemia Protein Dataset 2025
Using the DiseaseGrab scripts to extract the Viral Leukemia Protein Dataset from the Swiss-Prot viruses protein database. There are 103 proteins in the Viral Leukemia Protein Dataset. There were several errors that occurred on running the script.


021 Preparing the Databases for Formatting
Preparing the databases for formatting, downloading the diseaseMaesta scripts, and extracting these from the zip file. Renaming the databaseMaster file back to the databaseMaesta file to stop (my) confusion.
The databaseMaesta scripts are in a zipped file and will need to be extracted and placed correctly in the databaseMaesta folder before use (see video).

The databaseMaesta Scripts are:
Maesta.pl
fasta.pl
formatdb.pl (not to be confused with the formatdb.exe which is from NCBI)
databases.txt
Link to the databaseMaesta scripts on Patreon


022 Formatting All the Dat files with the Formatdb.exe and Creating the Fasta Files
Extracting the databaseMaesta zipped file again. Adding the sacoma dat file to the databases. Then running the formatting and fasta scripts via the Maesta.pl script. I forgot to mention the fasta files that are created but this is mentioned in the next video.


023 Looking at the Fasta Files
Looking at the fasta files that are created using the databaseMaesta scripts that I ran in the last video.


024 Disease Motifs Analysis of the Parkinson's Disease Protein Dataset
Using the Disease Motif scripts to do the proteomic analysis of the Parkinson's Disease Protein dataset.
Link to the DiseaseMotifs2025 scripts on Patreon

To see the protein sequence analysis on the proteins that are associated with Parkinson's disease (the Parkinson's Disease Protein Dataset) click on the following link:
Parkinson's disease dataset.

When you click on the "protein sequence analysis" link for various proteins on the subsequent pages, this will give you some idea of how the data outputs will look that are produced by the scripts contained within the books, although there have been some changes, so these will not be exactly the same. Some pages do have errors on them.


025 HGNC Human Genetic Information Download
The HGNC Human Genetic Information Download, and Running the GNgrab and GNXgrabDB scripts.

This file GNXgrabDB.pl needs to have the HGNC text file to be located in a directory where the path is ../HGNC/HGNC.txt. The data text will need to be downloaded from HUGO Gene Nomenclature Committee (HGNC) currently located at https://www.genenames.org. Key search terms are: HNGC, Statistics, Download.

You will also need to have downloaded the human protein database of protein datasheets which must be placed at ../databases/human.DAT. Download this database from Uniprot.


026 Setmaker
The Setmaker (TheSetmaker.pl) creates a series of sets that you can specify. In the Parkinson's Disease Dataset there are 73 sets that are created. Another programe (TheSetmakerLocation.pl) links the proteins with the loci of their genes.


027 nonPolyIndex
This creates the two webpages: the nonGIDS; and the Ranked nonGIDS.


028 Alzheimer's Disease
Running the Disease Motifs scripts on the Alzheimer's Disease Dataset. I had a bad cold when I recorded this video, so apologies for that, hopefully I did not make too many errors whilst I was explaining everything; but it is what it is.


029 GNgrab Changes to Partially Correct Errors
I noticed on the Alzheimer's Disease run, that there were a few small errors produced by the GNgrab perl script related to three proteins; SHMOS_HUMAN (doesn't have a GN name on the SwissProt datasheet), HUNIN_HUMAN, and BGIN_HUMAN (not in the HGNC file). I was trying to correct this but I feel that I have already spent a little too much of my time attempting to make changes to an overly complicated script, so I will leave this for a later time. It might be that when the external databases are updated these errors will disappear.


030 Clustal Omega
Using the online alignment programme Clustal Omega at the EBI to try align some proteins sequences. Used this on the following sets of proteins:

031 Alzheimer Subset (ShortSet)
Looking at the Alzheimer's subset which is the subset which is formed from incidences of the search term "alzheimer" in the FT part of the datasheets within the Alzheimer's Disease protein dataset. The disease motifs script were run on this set, which I labelled as a "short set" alzheimerShortSet.

I was also reminded of the terms "early-onset" (with hyphen) and "late-onset" (with hyphen) which I have added to the list of sets in both the Alzheimer's Disease and Parkinson's Disease protein datasets.

There are more recent videos on the Patreon site DiseaseMotifs711 Patreon site which will give you access to even more screen-capture-videos and Perl scripts which you may find of interest.



Useful Resources

Swiss-Prot protein databases can be downloaded from Uniprot

Swiss-Prot protein database of all proteins can be downloaded from Uniprot uniprot_sprot.dat.gz

Swiss-Prot taxonomic protein databases can be downloaded from Taxonomic Section at Uniprot

Download Strawberry Perl

Here you can download some programmes required (You might be better using Chrome as some browsers block these downloads) blastall.exe formatdb.exe
OR Download from NCBI site Index of /blast/executables/legacy.NOTSUPPORTED/2.2.19. For Windows (I think my scripts will only work on Windows) download the 32bit version blast-2.2.19-ia32-win32.exe

Download Padre

You might find this link to the SwissProt User Manual on the expasy site useful as it explains the annotation of the Swiss-Prot datasheets. Useful search terms to find this via a search engine are "Expasy", "Swiss-Prot", and "User Manual".

HUGO Gene Nomenclature Committee (HGNC). Key search terms are: HNGC, Statistics, Download.