Showing posts with label structure prediction. Show all posts
Showing posts with label structure prediction. Show all posts

June 20, 2011

Alternative side-chain structures from methyl CPMG

ResearchBlogging.orgAs I have mentioned before on this blog, the use of tools like CS-ROSETTA holds the promise of determining protein structures using only the chemical shifts of its backbone atoms. In addition to potentially making NOEs and RDCs redundant, this technology allows biologists to determine the conformations of minor members of the structural ensemble, which are very difficult to obtain using conventional approaches in population-dominated techniques like NMR and X-ray crystallography. There are two limitations here, however. First, we only gain insight into the backbone, and as we know, the positions of side chains in minor states can be critical for function. In addition, backbone chemical shifts are not always available due to relaxation problems. Both weaknesses could, in principle, be addressed by extracting conformational information from the chemical shifts of methyl groups, which report on side-chain behavior and continue to give good signal even in very large proteins. This is the rationale behind a series of recent papers from the Kay lab [1-3] intended to determine changes in side-chain rotameric state from methyl relaxation-dispersion data.


Read the rest...

March 10, 2008

Protein Structure from Chemical Shifts Alone

ResearchBlogging.orgUnless you have extremely good luck or a lot of supporting information, deriving a protein structure from NMR data is an enormous pain in the ass. First, you have to assign the resonances of the protein—that is, you must determine the chemical shifts of most or all of the protons in the protein, which in turn entails figuring out the chemical shifts of most of the carbon and nitrogen atoms as well. Then you have to acquire nuclear Overhauser effect (NOE) and/or residual dipolar coupling (RDC) data to figure out how far the atoms are from each other and how some of the bonds are oriented. Automated NOE assignment programs make the analysis of all this data less onerous than it once was, but particularly if your protein has non-ideal relaxation characteristics, scraping together enough data to derive a structure can be a tough task. The most irksome thing about it is that in principle, all the structural information you could ever want is contained in the chemical shifts you figured out in the very first step. What if you had a technique that could figure out a structure just from them?

First, a little more explanation about chemical shift. The local chemical structure dominates the chemical shift in most cases—you expect to find a proton in a particular place based on whether it's in a methyl group or bound to a nitrogen. Additionally, the chemical shift is sensitive to local bond angles. Surrounding groups (especially aromatic rings and paramagnetic atoms) can also alter the chemical shift substantially. However, the chemical shift is an ensemble averaged property. We cannot receive NMR data from a single molecule; as a result the observed chemical shift reflects every conformation in the ensemble and also (to some extent) the interconversions between those conformations. Because it should be possible to reconstruct an entire protein structure just knowing the local information about bond angles, it should in principle be possible to use chemical shifts to reconstruct the average conformation. The problem is that all these different factors get mashed up into a single number, often in contradictory ways. Parsing the purely structural factors (bond angles) out from this single number has proven difficult. Dihedral angle restraints based on chemical shifts have been used for many years, but only as a component of a more complete structural determination using NOEs or RDCs.

However, a series of publications over the past year or so has pointed towards some steady progress towards developing structures from chemical shifts alone, with a paper from the labs of Ad Bax and David Baker now in PNAS preprints showing some of the best progress yet (1). The approach used much resembles the CHESHIRE method described last year by Michele Vendruscolo (2), in that both are based on fragment replacement. The Vendruscolo group's paper explicitly compares CHESHIRE to David Baker's ROSETTA program. So it seems only natural for Shen et al. to incorporate refinements based on chemical shift directly into the ROSETTA program to create CS-ROSETTA.

The standard ROSETTA approach is to break the protein up into small overlapping fragments of several peptides. A library of structures (the PDB) is then searched with these fragments to obtain a set of about 200 potential conformations based on sequence similarity. ROSETTA then attempts to assemble low-energy (stable) structures out of these potential fragment conformations. CS-ROSETTA uses chemical shift data at two distinct steps. First, chemical shift data are used to select the most appropriate potential conformations from the library, theoretically improving the "building materials" for ROSETTA. In later stages, the consistency between the ROSETTA-predicted structures and the known chemical shifts is used to re-score their energy.

That this can significantly improve the ROSETTA output can be seen from the part of Shen et al.'s Figure 2 that I have shamelessly stolen for your benefit. These are predictions for calbindin (B) and HPr (C), with the ROSETTA predictions on top and the rescored energies on the bottom. As you can see, the calbindin structures do not have a well-defined energy minimum in the ROSETTA prediction, and the HPr structure has three minima which are not all close to the actual structure as measured by Cα root mean square deviations (RMSDs). Rescoring, however, produces funnel-shaped distributions of energy with respect to RMSD, such that low energies reliably indicate structures close to reality.

Shen et al. optimized CS-ROSETTA against 16 known structures. I checked their results back against the CHESHIRE results. Five proteins were predicted in both papers, and CS-ROSETTA did a better job in terms of backbone atom RMSD for four of them. On average, CS-ROSETTA produced a 24% reduction in this RMSD relative to CHESHIRE. Also, Shen et al. tested CS-ROSETTA blindly against nine proteins whose structures had been recently solved by the Northeast Structural Genomics Consortium, with favorable results.

This isn't the end of the road by a long shot. Backbone RMSDs for these predictions are generally <2 Å, which is easily good enough for picking out general characteristics of a fold. Identifying subtle features, however, will probably require higher precision and thus more rigorous refinement. However, having these predicted conformations in hand may significantly accelerate the assignment and refinement of structures using NOE data. Combining fragment-replacement approaches based on RDC data and chemical shift may also produce significant improvements.

There were other limitations. Shen et al. were not able to converge structures for every protein attempted. CS-ROSETTA is presently limited to proteins smaller than many routinely solved by NMR, and proteins with unusual or complicated topologies may not be solvable using this approach. And, of course, the presence of cofactors that significantly alter local chemical shifts will significantly complicate analyses of this kind, if not render them impossible. Obviously, a great deal of work remains to be done before computational approaches will be capable of tackling the large, highly degenerate systems where they would have the most power to resolve problems. However, the excellent results of CHESHIRE and CS-ROSETTA suggest that our ability to derive structures from limited NMR data will improve dramatically in the next few years.

1. Shen, Y., Lange, O., Delaglio, F., Rossi, P., Aramini, J.M., Liu, G., Eletsky, A., Wu, Y., Singarapu, K.K., Lemak, A., Ignatchenko, A., Arrowsmith, C.H., Szyperski, T., Montelione, G.T., Baker, D., Bax, A. (2008). Consistent blind protein structure generation from NMR chemical shift data. Proceedings of the National Academy of Sciences, 105 (12), 4685-4690. DOI: 10.1073/pnas.0800256105

2. Cavalli, A., Salvatella, X., Dobson, C.M., Vendruscolo, M. (2007). Protein structure determination from NMR chemical shifts. Proceedings of the National Academy of Sciences, 104(23), 9615-9620. DOI: 10.1073/pnas.0610313104

POSTSCRIPT: You can read another take on this paper at Plausible Accuracy.

Read the rest...

August 9, 2007

It came from the Protein Society! (Part One)

I went to the annual Protein Society symposium a few weeks ago and have finally had some time to organize my thoughts about it, so I figured I'd put some of them up here. Expect a couple of these to show up.

One of the most interesting presentations at the Protein Society (besides my own scintillating poster on field-cycling, haha) was Brian Volkman's poster on lymphotactin. This is a really interesting story that somehow seems to keep flying beneath the radar of most people, but intellectually it represents a giant challenge.

In a nutshell, the story is this: lymphotactin is a small signaling protein of the chemokine family, a group of proteins that are important for various kinds of regulation, including in inflammation and disease. Under fairly standard experimental conditions (200 mM NaCl, 10 °C) the protein adopts a normal chemokine fold, but at 45 °C (for reference, body temperature is 37 °C) and low salt, it takes on a totally different fold. You can read the original paper on this here. Of course, when you see something like this it's natural to ask what the physiological relevance of the finding is. Brian's poster at Protein Society basically answered this question by illustrating different biological roles for the two forms. The short version is that the conformational change appears to be some sort of regulatory switch. I'll have more on that in the next episode.

What I want to talk about here was what wasn't said about this at the meeting. After all, we were forced to witness the usual ninny-argument over whether folding was a linear pathway or a funnel of some kind. While it was refreshing to hear a lot more people pointing out that this distinction is more or less meaningless, it's odd that nobody is tackling the question through a protein like lymphotactin. Consider the following experiment: perform phi-value analysis or GdHCl-dependent HX experiments on lymphotactin to find what portions of the protein are structured in the folding transition state. We can imagine two outcomes.

In the first, the transition state is found to be completely different; that is, residues with high phi-values or the last residues to lose protection in the HX experiment are completely different for the two conformations. This would suggest that the latest common intermediate is the random coil (RC), and that the very first move towards a folded state dictates the state one finally arrives at (N1 or N2). Any intermediates (I1 and I2) along the pathway are unique to the end state, rather than shared between the two conformations (see right). This would fit most closely with the pathway view promulgated by Englander. Given that the hydrogen-bonding patterns are totally different for the two conformations, this might be expected.

On the other hand, it's possible that a residue or cluster of residues have similar phi-values, or lose their protection at a similar GdHCl concentration, between the two conformations. This would not be completely probative, as the similarities could quite easily be restricted to the observables and represent two different underlying structures. However, if veridical this might suggest that the latest common intermediate lies somewhere other than in the random coil. This would be more similar to the funnel view, in which a conformational search over the outcomes available to an intermediate gives rise to the ultimate choice of native structures. I've cartooned the idea over to the left. In this case I* represents a partially-folded intermediate that selects an end state based on the conditions.

After all, lymphotactin in both its native forms exists in physiological extracellular conditions. Knowing whether the protein must pay the full energetic cost of completely unfolding, or if it can switch conformations by taking a less-costly move to a common intermediate may be of significance to understanding the biology. And while these two alternatives (like the underlying models) are not as different as they may seem of first blush, the answer may do much to distinguish whether the conformational search of the funnel model or the deterministic folding of the pathway model is the best representation of the folding process.

Another major implication here is for the protein structure prediction crew. After all, the CASP-type experiments are all geared to the idea that a given sequence should give rise to a single folded structure. Lymphotactin is a clear counterexample to this idea, and while it may be unique, we certainly don't know nearly enough about proteins to let the dogma go unquestioned at this point. This makes for a much larger computational problem. I'm not a huge expert on these experiments, but my impression is that the programmers don't concern themselves too much about the characteristics of the solution the proteins are in. The assumption that co-solutes don't much matter is vastly simplifying, but as Brian's work shows, may ultimately limit the predictive power of these algorithms in significant ways.

Read the rest...