Monday, 28 September 2015

Pythonic spinner

Python is fun: it has lovely libraries, is a beauty to type and there are constant surprises — I only recently found out that 3.4 had introduced defaultdict() (collections library), which is phenomenal. With the web there are three options:

  • It can be used on the server-side on the web with the wsgi library or the Danjo framework. Open-shift is a great fremium script hosting service —I have used it here for example.
  • There are also some attempts to make JS parse python in the browser, namely Skupt and Brython, but as they convert the python code into JS code, so they are not amazingly fast (1), but you are showing the world python code.
  • One can transpile python to javascript, which is faster and less buggy, but that is unethical as you'd be serving JS and not a python script —CoffeeScript gets a lot of bad rep for that reason.

Thanks to CSS3 there are a lot of cool spinners out there to mark code that is loading, but none cater for python users. Therefore I made my own spinner icon, specifically: .
The code is hosted in my dropbox: https://github.com/matteoferla/Pyspinner.
<link href="https://rawgit.com/matteoferla/Pyspinner/master/pyspinner.css" rel="stylesheet"></link>
<span class=pyspinner></span>

(1) I tried Brython and liked that it had a mighty comprehensive series of libraries and that it had DOM interactions similar to JQuery. However, I could not get over the fact that, for me at least, changes to DOM elements were not committed until the code finished or crashed —which brought back bad Perl memories. Also the lack of CSS changes and the nightmare of binding functions to events makes me think I might try other options.

Publication-driven complexification

"Any sufficiently advanced technology is indistinguishable from magic."
—A. Clarke's Third Law
I like knowing what I am doing, it helps me figure out if something is wrong. However, there is a trend for accepted operations to become "mathemagical" either out of necessity... or out of publication race.  

RNASeq is often cheered for being absolute reads and thus requiring less normalisation than microarrays. Anyone that has tried to understand Loess or Lowess takes that as a blessing.
However, the situation is not that simple as different samples have different total amounts of DNA and, due to publications trying to out do the other it quickly gets complicated, where acceptant nodding is the best option lest one want to spend an evening reading a method's section and the appendix on a paper.
An arithmetic average is discouraged as the highly expressed genes will wreak havoc with other genes. On those lines, Mortazavi et al. 2008 (PMID: 18516045) introduced a per mille CDS-length normalised scale, which was called  RPKM —the letters don't mean much and feel like a weird backronym, which I don't get, although had I been them I would have made people have to use the ‰ glyph. The underlying logic is clear and for a while it was popular. RPKM got ousted by Anders and Huber 2010 (PMID:20979621) and Robinson et al. 2010 (PMID:19910308) with DESeq and edgeR, which do many calculations. The normalisation in DESeq is done by taking the median for each sample of each of the ratios of the gene count for that samples over the geometric mean of the values of that genes across the samples. That makes sense, albeit cumbersome to grasp. For the test of significance, the improvement on a negative binomial GLM spans a page or two and works by magic, or more correctly mathemagic, maths that is so advanced it might as well be magic —Google shows the word is taken by some weird mail order learn maths course or somesuch, but shhh! The sequel to DESeq, DESeq2 does everything automatically and in 3 lines everything is done. It is really good although I do like to know what I am doing.
Long story aside, I did in parallel the analysis with DESeq methods in MatLab, using their agonising tutorial, which was tedious, but good to see how it everything works  —to quote Richard Feynman: "what I cannot create I cannot understand". It wants a true waste of time as it made me realise what graphs would be helpful as the various protocols and co. try and out do each in terms of graph complexity. It really bugged me that an obvious graph, double-log plot of each replicate against its counterpart with a Pearson's ρ thrown in, was omitted everywhere in favour of graphs, such as the empirical cumulative function of the distribution of the variance against the χ-squared of the variance estimates —a great way of inspecting if there are any oddities in a distribution, but really not the simplest graph. It is used because it is sophisticated, which doesn't necessarily mean better (incorrect assumptions can go a long way in distorting data), but it means more publishable.
Curiously, once the alignment, normalisation and significance steps are done, it is cowboy territory especially for bacteria. For example, for bacterial operon composition there is either a program, Rockhopper, which does not like my genome (I had to submit for realignment each replicon on different runs to get operons for the plasmids and refuses to do the chromosome) and another paper where the supplementary zip file was labelled .docx and the docx was labelled zip, which goes to indicate the scarce interest.
The main use of differential expression is functional enrichment, which depends on annotations that are rough as nails...

I am in awe at the mathematics involved in the calculation, but I cannot help, but feel annoyed at the fact that they are embarrassingly more sophisticated than they ought to be, especially since the complexity gives somewhat marginal improvements and the created dataset will be badly mishandled with clustering based on wildy guessed functions.

Saturday, 12 September 2015

Unnatural amino acid biosynthesis

In the synthetic biology experiments with an expanded genetic code the biosynthesis of unnatural amino acid is not taken into consideration as the system is rather rickety and the amino acids unusual.

Similar amino acids

A lot of introduced amino acids are dramatically different from the standard set. However, there are several amino acids that never made it to the final version of the genetic code as they are subtly different from the canonical amino acids and must have been too hard for the high promiscuous primordial systems to differentiate. However, subtle differences would be useful for finetuning. Examples include aminobutyric acid (homoalanine), norvaline and norleucine, allo-threonine ("allonine"), ornithine or aminoadipate. If there was a way to introduce them and increase the fidelity it would be hugely beneficial. It would probably result in enzymes with higher fidelity and catalytic efficiency.
That Nature itself failed back then does not necessary mean that scientists would fail with modern metabolism. The main drawback is a selection system. Current approached to recoding rely on the new amino acid as fill in as opposed to something that makes the E. coli addicted to it. The latter would mean that the system could evolve to better handle the new amino acids. Phage with an unnatural amino acid have higher fitness (Hammerling et al., 2014). Unnatural RNA display (Josephson et al., 2005) could be used to generate a protein that is evolved to require the non-canonical amino acid as nearby residues are evolved to best suit that enzyme, if one really wanted to all that trouble. Alternatively and less reliably, GFP with different residues does behave differently and position 65 could handle stuff like homoalanine, but the properties between S65A or S65V are not too different. So if a good and simple selection method were present it could be doable.

Novel amino acid biosynthesis

This leaves with the biosynthesis of the novel amino acids, which is the main focus here. There are many possible amino acids to choose from, and a good source of information for that is the wikipedia article non-proteinogenenic amino acids  — I (reticently) wrote many years ago to sort out the mess that there was, but I subsequently left to the elements and it has become a bit cluttered like an unselected psuedogene. Why certain amino acids made it while other did not is discussed in a great paper by Weber and Miller in 1981.
Most of the amino acids that nature can make would just make structural variants. Furthermore, mechanistic diversity mostly comes from cofactors (metals, PLP, biotin, thiamine, MoCo, FeS clusters etc.), which is a more sensible solution given that there only one per certain type of enzyme. Some amino acids that are supplemented Nature cannot make with ease, in particular chlorination and fluorination reactions in Nature can be counted with one hand.

Homoserine, homocysteine, ornithine and aminoadipate.

E. coli already makes homoserine for methionine and threonine biosynthesis and homocysteine via homoserine for methionine. Ornithine is from arginine biosynthesis and was kicked out of the genetic code by it. Gram positive bacteria, which do not require diaminopimelate, make lysine via aminoadipate ("homoglutamate").

Homoalanine. 


The simplest novel amino acid is homoalanine (aminobutyrate). 2-ketobutyrate is produced during isoleucine biosynthesis, which if it were transaminated it would produce homoalanine. Therefore it is likely that the branched chain transaminase probably must go to some effort to not produce homoalanine (forbidden reaction). This amino acid is found in meteorites and is really simple, but its similarity to alanine, hence why it must have lost out.

Norvaline, norleucine and homonorleucine.

In the Weber and Miller paper the presence of branched chain amino acids was mentioned as potentially a result of the frozen accident. From a biochemical point of view, the synthesis of valine and isoleucine from pyruvate + pyruvate and ketobutyrate + pyruvate follows a simple pattern (decarboxylative aldol condensation, reduction, dehydration, transamination). The leucine branch is slightly different and is actually a duplication of some of the TCA cycle enzymes (Jensen 1976), specifcially it condenses ketoisovalerate (valine sans amine) and acetyl-CoA, dehydrates, rehydrates, reduces, decarboxylates and transaminates. If the latter route is used with ketobutyrate and acetyl-CoA one would get norvaline, with ketopentanoate (norvaline sans amine) and acetyl-CoA one would get norleucine. These biosynthetic pathway have actually been studied in the 80s. The interesting thing is that it can be used further making homonorleucine. Straight chain amino acids make sturdier protein, so there is a definite benefit there. The reason why norleucine is not in the genetic code is that it was evicted by methionine, which finds an additional use in SAM cofactor (Ferla and Patrick, 2014). Parenthetically, as a result the AUA isoleucine codon is unusal —it is also a rare codon (0.4%) in E. coli— as to avoid methionine has an unusual tRNA (ileX), which would make an easy target for recoding.

Allonine.

Threonine synthase is the sole determinant of the chirality of threonine's second centre.

Aminophenylanine.

Chorismate is rearranged to prephenate and then oxidatively decarboxylated and transaminated to make tyrosine, while the hydroxyl group of chorisate is swapped for amine for folate biosynthesis. If the product of the latter followed the tyrosine pathway one would aminophenylalanine.

Rethinking phenylalanine.

While on the topic, the whole route for phenylalanine biosynthesis is odd. It feels like an evolutionary remnant. If one were to draw up phenylalanine biosynthesis without knowing about the shikimate/chorismate pathway the solution would be different. If I were to design the phenylalanine pathway I would start with a tetraketide (terminal acetyl-CoA derived; 2,4,6,8-tetraoxononanoyl-CoA), cyclise (Aldol addition of C9 in enol form to C4 ketone), two rounds of reduction (6-oxo and 8-oxo) and three dehydrations (4,6,8-hydroxyl), followed by a transamination (2-oxo): phenylalanine by polyketide synthesis!

Naphthylalanine.


Phenylalanine has a single aromatic ring, naphthylalanine has two. Using the logic for the rethought phenylalanine synthesis we get a synthesis by hexaketide (2,4,6,8,10,12-hexaoxotridecanoyl-CoA). Namely cyclise (Aldol addition of C11 in enol form to C6 ketone and C13 to C4), two rounds of reduction (8-oxo and 10 or 12-oxo) and five dehydrations (4,6,8,10,12-hydroxyl), followed by a transamination (2-oxo). Naphthylalanine has been added to the genetic code (Wang et al., 2002), but was added as a supplement. GFP doesn’t work with either form of naphthylalanine (Kajihara et al., 2005), therefore there isn’t a good selection marker where the amino acid itself is beneficial. So probably the worst sketched pathway to possibly make.

Other aromatic compounds.

Secondary metabolites are often made by polyketide or isoprenoid biosynthesis, which are rather flexible so some extra compounds could be drawn up.

LEGO style.

All amino acid backbones are make in different ways and only cysteine, selenocysteine, homocysteine and tryptophan operate by a join-side-chain-on-with-backbone approach. Tryptophan synthase does this trick by aromatic electophilic substitution, where the electrophile is phosphopyridoxyl-dehydroalanine (serine on PLP after hydroxyl has left), while the sulfur/seleno amino acids are by the similar Micheal addition. So this trick could be extended to other aromatic, carbanions and enols, if one really wanted to, but that if far from a nice one size fits all approach.

Thursday, 9 July 2015

Thesis wordle

EDIT: The correct term is word cloud, while wordle is a website that runs on Flash and does not work in most browsers anymore. I would recommend Tagul instead. Here is the Wordle for my thesis:


I studied the enzyme MetC from Thermotoga maritima, which had alanine racemising activity in addition to a β-eliminating activity. I also studied the enzymes from Wolbachia, Pelagibacter ubique and E. coli. So far so go, all those words appear.
Unfortunately Wordle breaks up non breaking spaces, understores, interpuncts, except hyphens. It removes common words, but it does not collapse grammatical number, hence the enzymes/enzyme, genes/gene and activities/activity. In this version, hyphens were added and common plurals collapsed.
My appendix features a large Perl script, which affects the Wordle: else, elsif, foreach, print, file, sub and the name of a function (input). I assume "if" and "for" are weeded out by the filter against common words. Also YP and NP feature as there are many genbank accession identifiers.
If I remove the filter of common words the "the" is so prevalent that is squashes everything.

Now I want the raw data. So I pasted the thesis into an online word counter and saved the output and imported in MatLab and plotted the power-law distribution.




Not much else can be done with the data because many of the single-appearance words are actually sequences and there is no way to cluster the words by meaning. All the possible analyses will give generic results (e.g. more frequent words are shorter). If one were to go overboard, one option would be to track the progress of certain key words across the text. A Twitter trending equivalent. The problem is that I already know what words will appear more frequently were. So there isn't much point and it is probably best stopping at a Wordle step, which shows what one already knows, but is pretty.

Appendix, Scripts:
%first graph
loglog(x,'.');
ylabel('Counts');
xlabel('Rank of the unique word');
title('Word frequency in Matteo"s thesis');
offset= 300;
text(7,x(7)+offset,'MetC');
hold on;
plot(7,x(7),'r.','Markersize',20)
text(17,x(17)+offset,'enzyme')
plot(17,x(17),'r.','Markersize',20)
hold off;

>%second graph (not shown, frequent words are shorter).
x(isnan(x))=1;
xprime=unique(x);
z=cellfun(@length,words);
zprime=[];
yprime=[];
for i=1:numel(xprime)
    zprime=[zprime mean(z(x==xprime(i)))];
    yprime=[yprime numel(z(x==xprime(i)))];
end
plot(xprime,zprime,'ob');
hold on;
plot(xprime(yprime>1),zprime(yprime>1),'or');
hold off;
xlabel('number of counts of word');

ylabel('average number of letters for words of that frequency')

Sunday, 5 July 2015

Acknowledgement infographic

To break from tradition and avoid an "Oscars acceptance speech", on my thesis I made an acknowledge infographic. It was for fun and to acknowledge the folk who helped me keep my sanity during my PhD by procrastinating and to give a weighted credit to people, including Wayne's frowning.




Saturday, 4 July 2015

Getting the corresponding nucleotide sequences of protein sequences

Getting stuck on a really simple task is not a nice feeling.
A seemingly simple challenge I have faced a few times so far is getting the nucleotide sequence corresponding to a protein sequence: this seems really straightforward, yet it is not.

Wednesday, 6 May 2015

Speculations about methionine biosynthesis genes

Last year I wrote a review on the bacterial diversity methionine biosynthesis:
Ferla MP, Patrick WM. Bacterial methionine biosynthesis. Microbiology. 2014 Aug;160(Pt 8):1571-84.
PMID: 24939187, doi: 10.1099/mic.0.077826-0 and pdf.
It was crammed with facts and a couple of deductions that in my opinion are correct. However, there were a lot of hypotheses and conjectures, from plausible to wild, that did not make it into paper. Here I thought I might mention a few.

The MetCombo

It is my opinion that a bifunctional enzyme that catalyses both the MetC and the MetB reaction is impossible. I have come to call this hypothetical enzyme, the MetCombo. So the data at hand are:
  • MetC and MetB are close homologues and it is really hard to tell them apart in a phylogram —with the bold assumption that the uncharacterised genes are what have been guess.
  • Both KEGG and EcoCyc take the close homology to mean that bifunctional enzyme is present in several organisms —basically all those with MetB, which is a lot as you know from the met biosynthesis paper
  • They are in the same pathway
  • Nobody has ever seen a metCombo
  • Papers that try to evolve MetC ↔ MetB are not realy successful
  • Personal results: E. coli metC cannot rescue metB
  • Personal results: Thermotoga maritima "metB" is actually a metC and it has no in vivo or in vivo MetB activity(check out my thesis)
  • Catalytically a metC and metB in a single active site would be a disaster.

Catalytic profligacy

Cystathionine is a cysteine/alanine and a homocysteine/homoalanine joined together with a thioether. It has a short side (S is on the β) and a long side (S on the γ).
MetC is cysthationine β-lyase, it eliminates cystathionine at the thioether bond. On the shorter side (β).
MetB is cystathionine synthase it eliminates O-acetyl-homoserine at the ester bond and then attacks it with cysteine's thiol making cystathionine.
The two PLP enzymes hold cystathionine at some point but in radically different ways, one on the β side (MetC), the other on the γ (MetB).
Taking a step back, we have two types of cystathionine lyase and what controls the specificity between a β-lyase and a γ-lyase is not known —there have been a few papers looking into making MetC into a MetB and viceverse, but unfortunately nothing tackling this simpler issue. Cystathionine looks nearly identical from both sides: the sulfur bridge is hard to tell apart from a methyl group as there is only a slight size and charge difference. Methionine can be substituted with norleucine in protein with only minimal effect. Therefore it is intriguing how the enzymes bind it tightly in a specific way. My theory is that sulfur-π interactions may be involved as there are several tyrosines in the active site of MetC. Additionally, a β-elimination might be easier than a γ-elimination, therefore it is shame that there is a decent amount of data of the lack of γ-elimination activity in the β-lyase, but not viceversa. Therefore it would seem more likely to have a powerful bifunctional β- γ- cystathionine lyase than have retrained one that is strongly specific. However, this bifunctional enzyme is not an evolutionary a good idea due to the number of round trips it would do. Specifically, cysteine or homocysteine would go into making cystathionine, which the uncommitted lyase would either correctly transform or return a starting substrate —at the cost of ATP. The reason for this fascinating parenthesis is to conclude that cystathionine synthase/lyase combo that could do both, would be equally as bad of an idea —it would work due to flux from excess substrate to product in demand, but it is just extremely inefficient.

Conflicting results

Some methionine gene rescue experiments go in different ways that expected and there occasionally are concentration dependent oddities. My opinion is that this is due toone of the following:
  • It is dominating the threonine branch point and there is not enough threonine being produced.
  • It is depleting all the cysteine or homocysteine
I like the idea of enzymes being repressed from doing something easy, but bad on an aside (e.g. MetC can eliminate serine), but I don't think it's the case here. Unfortunately the only way to find out is to actually test whether the oddity perseveres when threonine and cysteine are added.

The ancestral PLP-dependent methionine gene

TBA

Methionine and norleucine

TBA

Alignment file

Here is the alignment file of manually aligned genes of various metB metC etc.