Search This Blog

Showing posts with label NCBI. Show all posts
Showing posts with label NCBI. Show all posts

Saturday, 27 February 2016

NCBI Sequence numbers for nitrile hydratase and nitrilase to 27/2.16

Looking at the bare search term "nitrile hydratase" amongst protein sequences (and remember, most aren’t but it’s a rough measure), today gives me 10224 hits, of which 4117 were RefSeq data. 
There are 80688 sequences labelled as "nitrilase" (not sure how robust that is currently), of which 20872 are pegged as RefSeq data.

Friday, 31 May 2013

NHase numbers to May 2013, and now nitrilase numbers too

Looking at the bare search term "nitrile hydratase" amongst protein sequences (and remember, most aren’t but it’s a rough measure), today gives me 4782 hits (+ 14% since last November), of which 2440 (+ 50% since August) were RefSeq data. It appears that there has been a lot of RefSeq going on with this enzyme class.
There are 24848 sequences labelled as "nitrilase" (not sure how robust that is currently), of which 12037 are pegged as RefSeq data.

Wednesday, 3 October 2012

New thermophilic NHase sequence

I have the NCBI database set up to send me a weekly email digest of newly uploaded nitrile hydratase sequences. One of the ones which appeared this week is from the proteome of Mycobacterium hassiacum DSM 44199. This organism is one of the Mycobacterium which use humanity as its primary habitat, and this particular one was isolated from a urine sample collected in Germany in 1995. The literature reports that this organism is comfortable with temperatures up to 65 degrees C, and can deal with up to 5% salt, all of which might well offer typical "moderately thermophilic" properties to its nitrile hydratase. The alpha chain shows a "CTLCSC" sequence suggesting it is cobalt-centred, and BLASTing this sequence shows that 7 other Mycobacterium species are between 89-85% sequence similar, with non-Mycobacterium sequences being significantly different (starting below 55% similar). Interestingly most of the other Mycobacteria which are sources of these sequences are from environmental samples with much lower temperature ranges. I cant find any reference to a paper where a Mycobacterium has been exploited for its nitrile hydratase activity either in the Prasad and Bhalla review from 2010 or from a Google Scholar search. Perhaps here's one to start with.

Monday, 24 September 2012

Diversity in Rhodococcus Sequences on NCBI

Downloading all the alpha chain amino acid sequences of nitrile hydratases of Rhodococcus origin from NCBI, you get over 100 sequences. By eliminating all the ones which are from PDB entries, you get 98 sequences. By using the usual amino acid sequence tag to identify which metal centre present, you get 68 iron-centred NHases and 24 cobalt-centred NHases, and six bits of rubbish. Of the iron centred ones, 61 have the VCSLC tag starting at position 109. None of the remaining have the metal binding region starting in the same place, and range from 96 to 149. There is much more of a spread with the cobalt centred sequences if you track the equivalent tag, as the histogram below shows.

Tuesday, 15 May 2012

How have the database numbers changed in a year?

Last year I did a quick text search for "nitrile hydratase" as a search term under proteins on the NCBI website. This gave me 2869, of which 1042 were RefSeq data. Today when I checked there are 3573 (+25%), of which 1369 (+31%) were RefSeq.
No new PDB files have been deposited of nitrile hydratases since March 2011 which was the NHase from Pseudomonas putida.

Friday, 23 September 2011

NCBI numbers for September

It has been a while since I had done the usual text search for "nitrile hydratase" under proteins on NCBI's website. Having done the proper filtering job on this numbers in August, I know there is a lot of duplication, error and bizarre search engine behaviour going on but I do think it shows how quickly the number of vaguely applicable sequences is growing. As of today there are 3066 hits for the target phrase (+63 since July), and 1125 RefSeq hits (+20 in the same period).

Disappointingly there was no change in the number of structures listed on the
Protein Data Bank.

Friday, 12 August 2011

Proper NCBI numbers for August

I have been looking for a number for the actual number of proper NHase alpha subunit sequences on record. I have done this by really combing through the NCBI protein records using a BLASTp approach for similarity, and excluding sequences that aren't an appropriate length nor have the usual active site binding motif. I have ignored environmental samples, and dropped out data from PDB files.
As of the start of this week (8th Aug 2011) I reckon there are about 190 NHase alpha sequences. That breaks down as about 25% iron and 75% cobalt centred. There are four from eukaryotic organisms (three marine and one plant), and the rest are... not.

Sunday, 12 June 2011

A better estimate of the number of NHases.

I have spent some time devising a spreadsheet into which you can download a list of FASTA data, and it will scan through looking for the "CTLCSC" fragment which indicates cobalt-centred and "CSLCSC" fragment which indicates iron centring, and total up how many hits there are of each in the list. It is all a bit more definite than just relying on "nitrile hydratase alpha" to give you accurate numbers!
Using the term "nitrile hydratase alpha" and then selecting the RefSeq selections, you get 121 cobalt centred NHases and 18 iron centred NHases. If on the other hand, you use all those sequences which are tagged as "bacteria", you get 278 cobalt-centred NHases and 131 iron-centred NHases. There are a few things to mention- these numbers don't include those sequences which have X in the tag to indicate that the cysteine is oxidized (these tend to be the PDB linked sequences) but do include some sequences which don't start with M and would make a molecular biologist wince, and the list hasn't been checked for duplication.

Tuesday, 17 May 2011

Searching with CTLCSC or CSLCSC

If you search the NCBI archives for nitrile hydratase enzymes just using that name you do get a large number of hits. Previous blog posts have recorded the size and rate of change of the number of hits you get if you do simple text searches with the phrase "nitrile hydratase". It is a slightly different story if you look at the RefSeq subset and then search through for the six amino acid sequence which is the metal binding motif for NHases. I am sure there are more elegant ways to do this but the way I did this was to download the relevant search with the sequences in FASTA format, load it up into MS Word and then use the "Find" function to look for each occurrence. [You can use the "reading highlight" tool and it gives you the count in the pop up box immediately].
Anyway the results were
CTLCSC (cobalt centred) -  123 hits
CSLCSC (iron centred) -  19 hits
So there are many more cobalt versions (87%) recorded than iron versions (13%) currently but given the incompleteness of the genome record, I reckon this may represent the current status of what that has been sequenced rather than the relative levels of occurrence in the wild.

Tuesday, 3 May 2011

NCBI numbers for May

A quick text search for "nitrile hydratase" under proteins on NCBI's website shows 2869 (+42 this month) hits for the phrase, and 1042 (+19 in the last month) RefSeq hits. It describes on the NCBI website how the RefSeq collection is their gold standard selection, so there has been high proportion of good stuff uploaded in the last month.

No change in the number of structures listed on the Protein Data Bank.

Thursday, 7 April 2011

NCBI- April numbers for NHase

A quick text search for "nitrile hydratase" under proteins on NCBI's website shows 2827 (+150 in a month!) hits for the phrase, and 1023 (+35 since the start of March) RefSeq hits. The number of PDB hits is up 5 to 41, as a new crystal structure from the cobalt centred NHase from Pseudomona putida has been uploaded along with the data for four point mutants of it.

Wednesday, 2 March 2011

NCBI- March numbers for NHase

A quick text search for "nitrile hydratase" under proteins on NCBI's website shows 2677 (+5 since the start of February) hits for the phrase, and 988 (+3 since the start of February) RefSeq hits. The number of PDB hits is unchanged at 36.

Wednesday, 2 February 2011

NCBI

Every so often I do a quick and dirty search for the text string "nitrile hydratase" on the NCBI website under the "proteins" database tab. Just now there were 2672 proteins annotated thus, of which 985 were RefSeq. Obviously this double counts the number of enzymes (at least) because NHase is two subunit enzyme and this search doesnt distinguish between alpha and beta chains, and there are always some misannotated things... but maybe there are 400-1000 different nitrile hydratase sequences out there?
A similar text search of PDB gives 36 structures. Some are pretty similar and are plus/minus something in the active site. A rough trawl suggests that's about a dozen distinct structures.