Wednesday, February 13, 2008
How do I submit homology models to FlexOracle?
FlexOracle, as well as hNMb and stonehinge, assumes the protein is solvated -- this is less than ideal for you since GPCR is a membrane protein. FlexOracle works by splitting the protein into two fragments at every possible pair of points, and computing the stability of the fragments using the FoldX force field. However FoldX will not guess at the effect of lipids, it can only estimate the effect of a solvent environment. That having been said, I have had some luck with membrane proteins, if only because often people are interested in some cytosolic domain. This may be the case for you.
TLSMD, on the other hand, uses the crystallographic B-factors and so it will tell you how it flexes in the crystal, which is often similar to how it will flex in vivo. All of these analyses are done for you when you submit to the HingeMaster server on molmovdb.org, as I encourage you to do. Bear in mind that the analysis will be done on one chain only. For particularly large proteins you may have to wait a couple of days. If you don't hear back from me when you get your results you are welcome to email me and I will help you interpret the results.
Threaded models have an additional strike against them in that the structure is almost certainly not properly equilibrated and thus is likely to be full of residual strains. This further limits the ability of FlexOracle to probe stability of fragments. hNMb is less sensitive to such details and so may be more useful, especially for the cytosolic domains. TLSMD will actually not return any results, since there will be no B-factors in a theoretical structure.
It still costs nothing to make a submission to our server, and despite these various issues you may still learn something.
Monday, February 11, 2008
How Do You Save or Download pdb file?
I need any pdb file like 119l or morph name: Small G-protein Arf6 with movie. I didn’t find an option to save or download and what software does it need to work on power point presentation to show protein motion.
A small movie is made on the old morph page. This is the easiest option but you can't control the perspective and have limited rendering options. More sophisticated movie making is available on our polyview3D page, administered by Alexey. Both are linked to from the top of the morph page.
Saturday, February 9, 2008
Resolving Conflicts in More than One Complex in the Same PFam Family
if there were two known complexes in the same pfam family, how did you use the combined information. And if homologs
were involved, did you just assign the aligned residues to the interface based on what interfacial residues were identified in the x-ray structure.
If we have known yeast complexes, this always takes precedent over any information in iPfam. If there is an interaction between two proteins known which share one Pfam domain _and_ the Pfam domain has been shown to bind itself in a crystal structure (i.e. in iPfam), we annotate that these two proteins will bind each other through the interface that is seen in that structure. In the
case of asymmetric binding between the same domain, the assignment is ad hoc. Ad II.) Yes, we use Pfam assignments as a form of homology mapping. We then assign the interface based on what is seen in the crystal structure.
Monday, January 21, 2008
PARE Perl Script
http://proteomics.gersteinlab.org/PARE_download/PARE.tar
Saturday, January 12, 2008
Missing information regarding mutation rates
I need to use your results for my research however, there are some missing information in your paper specifically in Figure2-A the following mutation rates are missing:
- G-> C with neighboring of C
- G->T with neighboring of C
- G->A with neighboring of C
We excluded CpG di-nucleotides from our analysis since they are known to have hyper mutation rates than any other di-nucleotides, due to the mechanism of methylation–deamination of cytosine. In fact mammalian genomes are depleted of CpGs, except of CG islands.
We also mentioned this in the paper.
Friday, January 4, 2008
Trouble Running Program to Reproduce Results in Defective Clique Paper
So I had some trouble reproducing the results you had in your paper. I got the negative and postive gold standard from this website: http://networks.gersteinlab.org/intint/supplementary.htm to pass in. The G+ matched up (8250), though by G- was off. These are the results I got running your large data set with the above mentioned gold standard: Initial G+ = 8250, G- = 2697594, bogus(?) = 7573, new edges = 388, initial max cliques = 4934, 61 new pos interactions, -37 neg interactions (total 98), and LR = -539.08. The results you got on your paper were as follows: G- = 2708622, new edges = 437, edges detected in gold standard = 73, in neg = 21; thus LR = 1141.3. So I definitely am passing in the wrong variables/datasets.
Can you please direct me to the datasets on gold standard you are using for the large dataset? I looked at your additional websites/supplements mentioned in your paper, but was unable to find it. If I am doing this correctly, can you then explain why I am getting different numbers? That would be very kind of you and greatly, greatly appreciated.
Furthermore, on running the 56by56 network, if I understand correctly, you just run it with out a negative gold standard? So all you are doing in your paper is comparing the maximum cliques created after clique completion, so I do not have to worry about LR results, right?
Thanks a lot for your interest in my paper. You are doing the right things. I got the same number as you did using the same gold standard sets. My only explanation is that this part were done by another co-author (Valery). I guess he must have slightly different sets. Unfortunately, I couldn't get in touch with him anymore (this is why the delay in my response). But, the numbers are in the ballpark. It does not in any way diminish the effectiveness of our method.
As for the 56x56 network, you don't need a negative set or worry about the LR results
Wednesday, January 2, 2008
Rice Genome Pseudogenes
Unfortunately, we have not run our PseudoPipe program on the rice genome. The PseudoPipe software itself is rather complex and designed to run on a computing cluster. (It is computationally intensive) If you wish to try to install the software anyway, the source is located at http://www.pseudogene.org/DOWNLOADS/pipeline_codes/