Showing posts with label proteins. Show all posts
Showing posts with label proteins. Show all posts

March 23, 2009

Rare codons give domains time to fold

ResearchBlogging.orgThe ribosome produces proteins by matching tRNA that has been correctly loaded with an amino acid to a codon (triplet of DNA bases) in the mRNA that contains the gene sequence. The triplet code allows 64 combinations of nucleotide bases, but proteins are made from only 20 amino acids (plus a "stop" signal). This means that most amino acids are coded by multiple codons, and hence have multiple tRNAs. Not all codons are created equal, however; in bacteria some codons are found much less frequently than others that represent the same amino acid. The tRNA associated with these "rare codons" is also less abundant than other tRNA, and this means that when a ribosome hits a rare codon, it often has to pause while it waits to encounter a loaded tRNA. To structural biologists like myself, who do their work by overexpressing proteins in bacteria, rare codons can be a nuisance because they slow down protein production, or even prevent it entirely. In a recent paper in Nature Structural & Molecular Biology, however, researchers from Germany suggest that the slowdown due to rare codons may have a functional advantage in vivo.

As a first step, Zhang et al. used a bioinformatics approach to survey the sequences of bacterial genes so that they could identify patches that would be slow to translate (the Methods section appears to contain an error in the description of this technique). They found that for proteins longer than about 300 amino acid residues, nearly every transcript contained at least one cluster of slow-translating codons. When the authors used a cell-free E. coli expression system to make some of these proteins and allowed only one round of translation initiation per ribosome, they saw a pattern of translation intermediates that matched the sizes predicted by the location of slow-translating patches.

In order to find out whether these translation intermediates had any significance, the authors examined the multi-domain protein SufI. In their prediction of the translation speed, which is on top in this figure that I have shamelessly stolen, there are four slow spots. Aside from the first one, these appear to correspond to the boundaries of different structural domains in the protein (lower part of the figure). Experiments with proteases suggested that these domains actually folded during the pauses, as the ribosome-bound translation intermediates were resistant to proteolysis.

Interestingly, when two rare leucine codons were replaced by more common ones (the authors call this SufIΔ25-28), the whole protein became vulnerable to degradation. Similarly, when extra tRNA for these rare codons was added to the cell-free expression system, the full-length protein became protease-sensitive. This suggests that the slow patches are actually necessary for proper folding of the protein. It's often the case that lowering the incubation temperature can improve the expression of certain proteins in E. coli. The authors of this study find that is also true for SufI, as the protease resistance of SufIΔ25-28 can be restored by lowering the temperature, and thus the overall translation rate. When analogous experiments with SufIΔ25-28 and tRNA supplementation were carried out in living E. coli, the translocation of SufI into the periplasmic space was reduced by a factor of 10 even though the overall protein concentration was not affected, indicating that the co-translational folding allowed by the rare codons is necessary for proper functioning of the protein in vivo.

Of course this is a single case study, and it would be premature to conclude that every patch of rare codons corresponds to an important co-translational folding event. Indeed, that doesn't even appear to be true of SufI, which folds properly when one of its other slow patches is removed. However, at certain key locations these stretches of rare codons may be an important part of the folding machinery in multidomain proteins. In addition, the more frequent appearance of rare codons in β-strands (as opposed to α-helices) may also be related to folding due to the slower kinetics of β-sheet formation. As the authors note, the intrinsic kinetics aren't everything — pauses in the translation process may also buy time for the complex to encounter essential chaperones or cofactors. Regardless of the mechanism, it appears that rare codons, in at least some instances, provide a way for the folding process to catch up with the translation process.

Zhang, G., Hubalewska, M., & Ignatova, Z. (2009). Transient ribosomal attenuation coordinates protein synthesis and co-translational folding Nature Structural & Molecular Biology, 16 (3), 274-280 DOI: 10.1038/nsmb.1554

Read the rest...

July 14, 2008

Microwave Pfu CelB 5 minutes for highest activity

ResearchBlogging.orgLike many bachelors, I regularly eat meals heated up in a microwave oven. I'd like to think I eat a somewhat lower percentage of frozen dinners than others in my situation, but even when it comes to food I cooked myself I usually don't have the time or patience to cook a recipe for one every night. That means I'm often eating reheated leftovers from the trusty Radar Range. The microwave oven works using dielectric heating, a process in which the movements of dipolar bonds are coupled to an oscillating external field. Water, for instance, is a molecule with a large dipole moment, and when you cook something in a microwave a lot of the heating occurs by the accelerated movements of water molecules. In principle this kind of excitation should also occur for other kinds of polar molecules, and thus we come to an interesting study from the lab of Alex Dieters, in which microwave radiation was used to activate a cellulase from a hyperthermophile.

Protein backbones consist of series of peptide bonds which include a carbonyl group, a classic example of a polar bond. Naturally, one might expect that the motion of these groups would be excited by microwave radiation. However, it does not directly follow that additional motion of the peptide backbone will actually accelerate chemical reactions, because this motion may be chaotic or unproductive. Moreover, from the fact that microwaves cook things (like eggs), we know that microwave radiation does a good job of denaturing proteins, sometimes at lower temperatures than we expect. Both of these problems can conceivably be avoided by studying a hyperthermophilic protein.

Proteins from hyperthermophiles such as Pyrococcus furiosus tend to be stable and optimally active at very high temperatures, at or even exceeding the boiling point of water. At lower temperatures, they retain their stability, but tend to become inactive. In many cases this reduction in activity appears to result from squelching internal motions that may be necessary to bind or properly orient substrates. Young et al. decided to study the β-glucosidase CelB from P. furiosus as a way of understanding whether microwaves might enhance enzymatic catalysis. Because CelB has optimal activity at 110° C it should be possible to see a significant difference in activity if microwave activation works. The stability of this protein at high temperatures also suggests that you will not accidentally cook it.

Sharp readers will have noticed an obvious problem with this idea—because heat activates this protein, and microwaves heat aqueous solutions, we must incorporate some kind of control in order to determine the pure effect of the radiation as opposed to the temperature. Young et al. resolve this problem by monitoring the heating of the sample during microwave irradiation, and then using a normal thermal apparatus to match this temperature profile (Figure 2). When a reaction reached 40° C using either heating method, it was quenched by the addition of a basic solution and the concentration of products was measured. Simply heating the Pfu CelB reaction to 40° C produced negligible activity, but microwaving it increased the activity by 4 orders of magnitude (i.e. a factor of 10,000). Less dramatic, but still significant, effects were observed for two other hyperthermophilic enzymes, but an enzyme from a mesophilic organism (the almond) was not activated by microwave irradiation.

That the microwaves caused increased backbone motion was supported by the finding that irradiating Pfu CelB at 75° C caused it to denature; this temperature is well below the normal melting temperature of this enzyme (115° C). The authors attribute the differences in activation between CelB and the other hyperthermophiles to their lower optimal activity temperatures, but it is also possible that the particular motions enhanced by microwaves are simply not as productive in those molecules. Although all the dipoles should be affected in similar ways by microwaves, they are all oriented differently with respect to each other in the protein molecule. As a result, the induced motion may be chaotic, perhaps specifically so, and therefore the activation of a particular thermophile may depend on the nature of the motions needed for its catalytic cycle. Enzymes that require large ensemble motions of subdomains, such as adenylate kinase, might not be activated as much as a protein that simply needs to be melted a little. Examining the differences in structural dynamics of enzymes differentially activated by microwaves may be an interesting area of future study.

While microwave activation is unlikely to revolutionize some of the more common uses of hyperthermophilic proteins (i.e. PCR), it does have promise. In ligations, for instance, hyperthermophilic enzymes can not be used at present because many DNA inserts denature at the optimal active temperature. With microwave activation, it may be possible to employ extremophilic ligases in these reactions, gaining the benefits of their speed and durability without having to worry about accidentally melting your DNA. Depending on the enzymes available, this technique may also prove valuable in improving mobile medical laboratories and developing novel diagnostic tools for field work.

1. Young, D.D., Nichols, J., Kelly, R.M., Deiters, A. (2008). Microwave Activation of Enzymatic Catalysis. Journal of the American Chemical Society DOI: 10.1021/ja802404g

Read the rest...

June 23, 2008

A conformational equilibrium controls the Vav DH domain

ResearchBlogging.orgOne emerging view of allostery, and protein behavior generally, describes function in terms of pre-existing equilibria. In this view, proteins are not like switches that get turned on and off, but rather are like dials that are turned "more on" or "more off" depending on the conditions. In this view, regulatory modifications such as phosphorylation do not enforce an active conformation so much as promote it. Because some relaxation measurements are sensitive to conformational exchange, NMR is well-suited to examine systems with this behavior. In the most recent edition of Nature Structural and Molecular Biology, a group from Michael Rosen's lab discover that this kind of equilibrium is governing the behavior of the DH domain from the Vav protein.

Vav activates Rho GTPases by inducing the exchange of GDP for GTP, making it a Guanine nucleotide Exchange Factor (GEF). This activity is performed by the DH domain, and is inhibited by a small neighboring element called the acidic region (Ac). As you can see from the figure of the combined Ac and DH domains (AD) at right, Ac (red) inhibits the DH domain (blue) by forming a helix that binds to the active site (explore this structure at the PDB, noting that the numbering is off by 167). Phosphorylation of Y174 (red side chain) unfolds the helix and exposes the active site, which would be a nice model except for two things. First, as you can see, Y174 is pretty well buried in this structure, which would make it difficult to phosphorylate. Second, mutation of Y174 to phenylalanine, a residue that cannot be phosphorylated, activates DH domain activity. How can Y174 get phosphorylated? What is phosphorylation actually doing?

One possible answer is that the existing structure doesn't tell the whole story. A protein, after all, doesn't have just a single structure, but rather an ensemble of structures across a population or time. While AD may spend most most of its time in this inhibited state, it's possible that sometimes it adopts an alternate conformation that allows Y174 phosphorylation. Li et al. set out to assess this possibility using measurements of R2 relaxation dispersion. Residues that have a large field-dependence of transverse relaxation (ΔR2) are undergoing some sort of conformational exchange process that changes their chemical shift between two or more states.

Using CPMG experiments on methyl-bearing side chains, Li et al. identified two groups of residues engaged in conformational exchange processes, shown at left. The first group (orange side chains) have a large ΔR2 that vanishes once Y174 is phosphorylated. The second group (green side chains) have high ΔR2 in both the phosphorylated and unphosphorylated states, but the rates are slightly higher in the former. Residues that were observed to have low ΔR2 are shown with gray side chains. All of this indicates that there are two dynamic processes occuring on the microsecond-millisecond timescale. The first encompasses some change in the chemical shift of the acidic helix, while the second involves some unknown process. However, because the Group 2 residues react to the phosphorylation state, these processes are likely linked in some way. I notice that the Group 2 residues are clustered around loops and joints in the upper half of the domain (in this view), while residues not adjacent to loops or joints do not appear to have significant ΔR2. It is possible that the observed dispersion represents some flexing of the domain around these loops, and that the rate of this motion increases slightly when the binding site is unoccupied.

That's all very interesting, but it's also bad news for the analysis, because it means the observed relaxation dispersion would have to be fit to a four-state model in order to obtain populations and kinetic parameters. Previous analysis by several groups has shown this to be a dubious proposition, so Li et al. take an alternative approach. Rather than try to fit out populations from the dispersion data, they make a series of mutations to AD to push the populations of the two states in one direction or the other. They find a number of states where the methyl peaks lie on a line between the open (phosphorylated) and bound (unmodified) states. The Y174F mutation lies very close to the phosphorylated state, interestingly enough, implying that the phosphate group itself is not a significant determinant of chemical shifts in the open state. Using a combination of HSQC peak positions and ΔR2 measurements, Li et al. determine for each mutant or modification what population of the ensemble is in the open state. They find that this NMR-assigned population correlates with the rate constant (kcat/KM) for phosphorylation.

This implies a model in which regulation of DH by Ac involves an equilibrium between the bound and open states. In the bound state, Ac forms a helix in the binding site, an effect strongly dependent on a hydrogen bond to the OH of the Y174 (R332 may be the partner here). However, in this state Ac samples the open state about 10% of the time. While in the unbound state, Ac can be phosphorylated, a modification that prevents helix formation or binding; probably by steric interference in the binding site. This stabilization of the open state dramatically increases the chances that the Vav DH domain will be in an active state when it encounters a target. Thus, DH regulation depends on a population shift of an underlying equilibrium, not a singular on/off switch. This model has the advantage of accounting for how Y174 gets phosphorylated and why a mutation that prevents phosphorylation nonetheless leads to a constitutively activated state.

In vivo, Vav consists of many other domains in addition to the AD construct used here. These domains are known to contribute to the inhibition of DH; given these results it is probable that they do so by stabilizing the bound state. Because Ac binding causes the formation of a negatively-charged surface on one side of the helix, charge stabilization is a likely mechanism. Further research will hopefully identify these mechanisms, as well as the origin of the second conformational exchange process revealed by the experiments in this paper. This study is a good example of scientists making the best use of limited data to describe an instance of this important, but difficult to characterize, regulatory mechanism.

1. Li, P., Martins, I.R., Amarasinghe, G.K., Rosen, M.K. (2008). Internal dynamics control activation and activity of the autoinhibited Vav DH domain. Nature Structural & Molecular Biology, 15(6), 613-618. DOI: 10.1038/nsmb.1428

Read the rest...

June 14, 2008

NSAIDs bind to amyloid-β

ResearchBlogging.orgOne of the best-known features of Alzheimer's disease pathology is the formation of proteinaceous amyloid plaques in the brain. In Alzheimer's disease these plaques are primarily formed by the amyloid-β peptide (Aβ) derived from the amyloid precursor protein (APP) by the action of β- and γ-secretase. The length of the Aβ peptide varies, but the 42-residue form (Aβ42) is more likely to form plaques and fibrils. Although it remains uncertain whether plaques are a cause of Alzheimer's disease symptoms, or merely an effect of some underlying derangement, finding some way to prevent or reduce plaque formation is a major goal in the field. This week in Nature, a team of researchers from institutions all over the US and Europe show that non-steroidal anti-inflammatory drugs (NSAIDs) may be able to accomplish these goals by binding to APP and Aβ directly.

Previous research from the Koo lab indicated that some NSAIDs specifically reduced the production of the amyloidogenic Aβ42 fragment (1) both in cultured cells and in a mouse model of the disease. APP was still processed into peptides, but these were shorter and less likely to form amyloid plaques than Aβ42. Significantly, the cleavage of other γ-secretase targets was not affected, meaning that side-effects of NSAID treatment might be minimal. Although NSAIDs were expected to ameliorate Alzheimer's symptoms by reducing inflammation, Weggen et al. found that the beneficial effects were not the result of cyclooxygenase inhibition. In a follow-up paper (2), Weggen et al. used experiments on cultured cells to show that the drugs were directly modulating γ-secretase activity. These experiments also showed that mutations to presenilin-1, a core component of the γ-secretase complex, could either increase or decrease the effect of NSAIDs, suggesting that it was the protein directly affected by these drugs.

Kukar et al. set out to test this hypothesis using photaffinity labeling. They took a few compounds known to alter Aβ42 levels and added a functional group that would react with a protein in the presence of UV light. These covalently-labeled proteins could then be detected, and this would serve as a relatively easy way to determine which component of the γ-secretase complex was actually binding NSAIDs. Like many cleverly-designed experiments, this failed in an interesting way: no known components of the γ-secretase complex were labeled. Fortunately, the researchers realized that there was another component to the complex they hadn't tested yet: the substrate.

It turned out that the NSAIDs could label a 99-residue fragment of APP. Moreover, this labeling was reduced by other γ-secretase modulators (GSMs) and unaffected by non-GSM NSAIDs. Using a series of progressively shorter constructs, Kukar et al. localized the binding activity of GSMs to residues 28-36 of amyloid-β.

This on its own is a very useful finding because it provides a target for refinement of these compounds. Knowing where and to what protein a possible drug binds makes it easier to develop assays to test new potential drugs, as well as enabling structure-based design. However, the authors took the next step and asked whether these drugs, because they bind to a region of APP known to be involved in the formation of amyloid plaques, might inhibit plaque formation directly. In cultured cells, they found that treatment with certain substrate-targeting GSMs decreased the formation of Aβ dimers and trimers even under conditions where the overall concentration of Aβ42 was not altered.

This suggests that these GSMs may be able to fight the buildup of amyloid plaques in two ways. By altering where γ-secretase cleaves APP, they reduce the concentration of Aβ42. Moreover, by interfering with Aβ oligomerization they fight the formation of plaques directly. With luck, further work in medicinal chemistry will arrive at compounds that enhance both these activities. The development of compounds that significantly reduce or prevent the formation of amyloid plaques will be a great step forward for Alzheimer's research. Even if such drugs do not prove to be a cure, a clear indication that plaques don't cause Alzheimer's would be a critical insight.

I want to emphasize that although these results are quite promising, they do not prove the efficacy of NSAIDs in ameliorating actual Alzheimer's symptoms. Transforming these findings into a cure or even an effective treatment will require a great deal of additional research, if it is even possible. You should not attempt to treat Alzheimer's with NSAIDs, or begin a regimen of NSAIDs or any other kind of drug or supplement, unless you have first discussed the possible risks and benefits with your doctor. And no, Minnesota, I do not mean a naturopath.

1. Weggen, S., Eriksen, J.L., Das, P., Sagi, S.A., Wang, R., Pietrzik, C.U., Findlay, K.A., Smith, T.E., Murphy, M.P., Bulter, T., Kang, D.E., Marquez-Sterling, N., Golde, T.E., Koo, E.H. (2001). A subset of NSAIDs lower amyloidogenic Aβ42 independently of cyclooxygenase activity. Nature, 414(6860), 212-216. DOI: 10.1038/35102591

2. Weggen, S. (2003). Evidence That Nonsteroidal Anti-inflammatory Drugs Decrease Amyloid β42 Production by Direct Modulation of γ-Secretase Activity. Journal of Biological Chemistry, 278(34), 31831-31837. DOI: 10.1074/jbc.M303592200 OPEN ACCESS

3. Kukar, T.L., Ladd, T.B., Bann, M.A., Fraering, P.C., Narlawar, R., Maharvi, G.M., Healy, B., Chapman, R., Welzel, A.T., Price, R.W., Moore, B., Rangachari, V., Cusack, B., Eriksen, J., Jansen-West, K., Verbeeck, C., Yager, D., Eckman, C., Ye, W., Sagi, S., Cottrell, B.A., Torpey, J., Rosenberry, T.L., Fauq, A., Wolfe, M.S., Schmidt, B., Walsh, D.M., Koo, E.H., Golde, T.E. (2008). Substrate-targeting γ-secretase modulators. Nature, 453(7197), 925-929. DOI: 10.1038/nature07055

Read the rest...

June 13, 2008

DegP 24-mers are spacious chaperones

ResearchBlogging.orgA few weeks back I mentioned a paper in PNAS on allosteric regulation of DegP protease function by its PDZ domains. This week in Nature, the same group (same first author, even) provides some intriguing new insights into the workings of this combined chaperone and protease. Using different arrangements of a base trimer structure, DegP assembles into multimers of 6, 12, and 24 protein units. Krojer et al. determine the structures of the larger complexes using cryo-electron microscopy and X-ray crystallography, and how these structures might achieve the refolding and degradative functions of this HtrA protein family member.

As I mentioned last time, a single DegP protein consists of a protease domain and two PDZ domains, a domain I talk about a lot. This makes for a decently-sized protein, but in fact DegP is rarely encountered in vivo as a monomer. It is known to form trimers and hexamers, and now dodecamers and whatever fancy Latin or Greek word you would use for a 24-mer. In all of these cases what we're really dealing with are higher assemblies of trimers. For instance, the dodecamer is a tetramer of trimers. In addition to its ability to degrade proteins, DegP is known to have a chaperone function, and also to be able to shepherd OMPs (outer membrane proteins) through the periplasm of E. Coli.

In the case of the 24-mer, the DegP molecules assemble into a giant, hollow octahedral shape with an interior cavity 110 Å wide, which is larger than the cavity of the well-known chaperone GroEL. The PDZ domains mediate contact between adjacent trimers, and the protease domains form the "faces" of the octahedron. From the first figure in the paper it almost looks like you could cram 2 folded OMP proteins into the cavity formed by this structure. This oligomeric complex is so large it could conceivably span the whole periplasmic space between the inner and outer membranes of an E. Coli. Because the 24-mer has fairly large pores, it seems possible that it could form a tunnel that protects OMPs from aggregation and degradation as they cross the periplasm. In addition, positively-charged residues concentrated on the surfaces of the PDZ domains appear to give the multimer some affinity for membranes; these positive charges are concentrated around the edges of the pores.

To check this idea in vivo, Krojer et al. made a DegP-null strain of bacteria and measured the concentration and location of OMPs. They found that for several OMPs, deleting DegP did not change the concentration of the OMP in a whole cell lysate, but reduced levels of these proteins in the outer membrane. Further experiments indicated that DegP oligomers can protect OMPs from proteases. DegP itself can degrade unfolded OMPs, but stabilizes the folded proteins.

The researchers also managed to catch a glimpse of an OMP inside a DegP oligomer. In the case of the 24-mer this was difficult, probably because the sheer size of the enclosed space allows so many orientations of the OMP that getting a regular structure is impossible. However, they found that the structure of DegP dodecamers bound to OMP was fairly homogeneous, allowing an investigation by cryo-EM. You can see an image of the structure (shamelessly stolen from Figure 5) at left; the DegP molecules are in warm tones, and a molecule of OmpC is visible in blue. Again, the protease domains form the faces of this tetrahedral cage, while the PDZ1 domains (not PDZ2 in this case) mediate trimer-trimer contacts.

But what about DegP's protease function? Chromatography experiments (Figure 2) suggest that at room temperature, the presence of unfolded substrate molecules induces the formation of higher-order oligomers. These experiments cannot tell us whether these 12- and 24-mers have the same conformations as determined in the experiments above, but the fact that the PDZ1 domains mediate trimer-trimer contacts in both complexes suggests a possible mechanism for the allosteric activation noted previously. However, at higher temperatures where protease activity is markedly increased, the dominant species in solution was the trimer itself, even in the presence of substrates.

The authors propose that the hexameric form of DegP is a resting state, and that DegP12 and DegP24 form in response to specific stimuli. This model certainly fits the observations in solution, but it seems possible to me that the presence of membranes could induce formation of DegP24. In vivo studies using fluorescence or TEM may be able to address what form actually predominates in the periplasm. Additionally, these results do not directly address the role of oligomerization in the switch between proteolytic and chaperone function. Particularly crucial in this regard is the question of whether the larger complexes that form during proteolysis are the same as those that form around OMPs. While it's reasonable to think that they are, the demonstrated structural versatility of DegP trimers suggests that these large assemblies may represent alternative conformations. Alternatively, it could be that features of the periplasm make DegP24 and DegP12 poor proteases in bacteria, and that the more rapidly diffusing naked trimer is the only efficient protease in vivo among DegP oligomers.

If, however, the chaperone and protease oligomers are structurally equivalent, a whole new class of questions opens up. Is formation of 12- and 24-mers sufficient to activate proteolysis, or is some additional step required? What protects folded OMPs from degradation by the protease subunits? These are challenging questions, but the new structures will be of great assistance in guiding the design of the genetic and biochemical experiments that answer them.

1. Krojer, T., Sawa, J., Schäfer, E., Saibil, H.R., Ehrmann, M., Clausen, T. (2008). Structural basis for the regulated protease and chaperone function of DegP. Nature, 453(7197), 885-890. DOI: 10.1038/nature07004

Read the rest...

May 29, 2008

PDZ domains allosterically regulate the bacterial envelope protease DegP

ResearchBlogging.orgAnyone who reads this blog regularly knows that I have a great deal of interest in the allosteric potential of the PDZ domain, a small protein binding domain that can be found in every branch of the tree of life. In previous posts I've discussed the evidence for allosteric communication within the PDZ domain, as well as evidence for long-range energetic interactions. An upcoming paper emerging from a collaboration of several European groups provides yet another example of the PDZ domain's allosteric potential bearing fruit, in this case in the regulation of a protease. The article is open-access, so go ahead and open it up in another window to follow along.

Krojer et al. are investigating the heat response in bacteria. Just as your body responds to ambient warmth (by sweating, etc.), bacteria have several systems to help them cope with high temperatures. These include systems for folding proteins that have lost their shape due to the heat, and also systems that degrade proteins when the refolding system can't keep up. A protein that breaks down other proteins is called a protease, and enzymes of this type are widely used in nature (blood clotting, viral maturation, digestion, etc.). Because of their destructive potential proteases are often tightly regulated, as is the case with the protease in this article, DegP. In fact, DegP can also serve as a refolding protein (or chaperone)—some regulatory mechanism causes a switch between these functions. DegP is an E. coli enzyme, but has a similar architecture to some human enzymes linked to diseases that involve protein misfolding and aggregation.

Krojer et al. performed experiments to characterize the proteolytic activity of DegP, mostly summarized in Figure 1. These results indicate that DegP is processive, i.e. that it makes many cuts on a target rather than just one. Using mass spectrometry the researchers determined that the targets got cut up into chunks 8-22 residues long, with the most common lengths being 12 and 17 residues. The protease preferred to cut proteins after a valine, alanine, isoleucine, or threonine: these are very common residues in proteins and interestingly most are β-branched. However, when they made short peptides that matched the apparent cleavage pattern, they did not observe any proteolysis.

From this the authors concluded that binding to the PDZ1 domain of DegP was necessary for the proteins to get cut. They performed a series of experiments with longer peptides that showed that a proper binding site 13-17 residues away from the cleavage site was necessary to get proteolysis. Also, the C-terminal residues preferred by PDZ1 are the same as the ones where the protease cleaves. Modeling an unstructured peptide substrate into the known structure of the DegP complex (Figure 4) indicates that a peptide chain ~16 residues long is needed to reach from the PDZ1 domain of one DegP to its own protease domain, and a chain ~12 residues long is needed to stretch to the protease site of an adjacent DegP molecule.

Well, this suggests a tidy little model. The C-terminus of an unfolded protein binds at the PDZ domain and is cut by the protease about 12 or 16 residues down the line. The cut produces a new C-terminus, which binds at the PDZ domain, and the process repeats. In this way the DegP complex processively degrades unfolded proteins and the bacteria are saved from toxic aggregates. But this is not the whole story. You see, it turns out that DegP is pretty efficient at cutting peptides even if the PDZ substrate and the cleavage site are not attached to each other.

The above model suggests two predictions. First, a peptide that binds to the PDZ1 domain (ALE peptide) but cannot be cleaved by the protease should inhibit the proteolysis of an unfolded protein. Second, the activity of DegP towards a substrate that is too short should not be affected by adding an ALE peptide. However, when the researchers in this case performed these experiments the results were quite different than expected. The ALE peptide slightly activated the degradation of an unfolded protein, and enhanced the cleavage of the short peptide 50-fold. Further experiments indicated that the binding of a peptide to the PDZ1 domain of one DegP molecule activated the protease of that molecule and one neighboring molecule of the complex.

So now the model is more complex, but also much more interesting. Binding of an unfolded protein's C-terminus to the PDZ1 domain does help generate the processivity and molecular ruler effects mentioned previously. But this binding also allosterically increases the intrinsic rate of proteolysis in nearby protease subunits. The effect is a sort of positive feedback mechanism that keeps the protease running until the whole target protein is degraded. This adaptation means that the protease can activate rapidly when there are a lot of free termini floating around, but won't go chopping up every flexible loop it comes across.

The precise mechanism by which binding at the PDZ site activates the protease is beyond the scope of the present study. It may be that the networks suggested by previous research do not come into play in this instance. Looking at the biological complexes predicted from the crystal structure (explore it at the PDB) I would guess that binding-dependent remodeling of the strand linking the protease domain to PDZ1 may be responsible, rather than the dynamic networks. Future experiments will probably address this allosteric mechanism.

1. Krojer, T., Pangerl, K., Kurt, J., Sawa, J., Stingl, C., Mechtler, K., Huber, R., Ehrmann, M., Clausen, T. (2008). Interplay of PDZ and protease domain of DegP ensures efficient elimination of misfolded proteins. Proceedings of the National Academy of Sciences 105(22), 7702-7707 DOI: 10.1073/pnas.0803392105 OPEN ACCESS

Read the rest...

March 8, 2008

Why are Xfaso and Pfl Cro so different?

A few weeks ago, when I posted on the transitive homology studies performed by the Cordes group, I promised a closer look at the structures when they became available. If you'll recall, one of the central findings of the Roessler et al. paper (1) was that the Xfaso and Pfl 6 Cro proteins, though they had 40% sequence identity, as well as an identical function, had very different structures and dimerization characteristics. The Pfl6 Cro structure is now available in the PDB, and Dr. Cordes was kind enough to send me the Xfaso Cro structure, which has been held up by some technicalities. I made the overlay of the structures to the left, with Xfaso in dark green and Pfl 6 in crimson. As you can see, the N-terminal helix-turn-helix motifs of the two molecules overlay very precisely, with some slight differences in orientation in the context helices. The C-terminal portions, of course, are completely different. How did they get to be this way?

Well, I have a few thoughts. To achieve a significant change in structure like we have here, two possibilities suggest themselves. We can destablilize one structure, or we can stabilize the other. So let's try to look at this from both angles. First, the Xfaso structure, which is on the left. The ribbon is orange for residues that don't change between the two proteins, green for mutated sites, and red over a deleted range. I've also drawn in a couple of the mutated side chains that might have an effect. For instance, at the upper right you can see a glutamate of Xfaso Cro that becomes a glycine in Pfl cro. The presence of the glycine may destabilize the helix. Lower on the helix, a solvent-exposed arginine becomes a hydrophobic leucine in Pfl; likely the structure will change to reduce the contact of the leucine with water. By the same token, a partially-buried threonine at the base of the helix gets mutated to glutamate. Not only might this put an unsolvated negative charge inside a hydrophobic region, but favorable helix-capping interactions of the threonine might be broken. Roessler et al. also point out that a pair of cysteines in the Xfaso structure are in a favorable position to form a disulfide bond; both are absent from the Pfl6 sequence.

On the other side of things, what mutations are stabilizing the new fold of Pfl 6 Cro? Check out the image below, color-coded the same way as the previous structure:
There are a couple of things going on here. First, the hydrophobic residues. The leucine that was an arginine in Xfaso Cro has indeed been buried, up at the top of the structure. Nearby, the glutamate that was a partially-buried threonine in Xfaso is now fully solvent-exposed. In the same vein, an aspartate has become an isoleucine near the middle. This new isoleucine appears to pack into the core of the companion protein near a conserved isoleucine (and a new methionine), and probably accounts to some degree for the stability of the dimer. There are a wealth of minor effects too—an alanine to arginine mutation has created a charged group that is solvent stabilized on one of the strands; down near the bottom a proline-to-arginine change probably has allowed the extension of a helix. Some other side-chains are drawn that I was taking a look at, but probably aren't critical, and of course I've probably missed a few that matter.

The R→L, T→E, and D→I mutations probably do a great deal to change the most stable conformation from all-α to α/β and strengthen dimerization in solution. However, just from looking you can never be sure what mutations are the most important. Doubtless the Cordes lab is already examining which residues are critical in altering the conformation. It will be interesting to see whether point mutations tend to move Xfaso from the all-α to the mixed structure, or whether they produce molten globules.

1. Roessler, C.G., Hall, B.M., Anderson, W.J., Ingram, W.M., Roberts, S.A., Montfort, W.R., Cordes, M.H. (2008). Transitive homology-guided structural studies lead to discovery of Cro proteins with 40% sequence identity but different folds. Proceedings of the National Academy of Sciences, 105(7), 2343-2348. DOI: 10.1073/pnas.0711589105

Read the rest...

March 5, 2008

Allostery without conformational change

ResearchBlogging.orgAllostery is a strange-looking word for a relatively simple idea: regulation at a distance. Binding events at one location on a protein can influence binding events that are relatively far away. It is allostery—in the form of cooperative binding in hemoglobin—that makes our oxygen-delivery system work. Because it provides an alternative way to attack drug targets, allosteric regulation is an attractive possibility for new medicines; recent results suggest that allosteric inhibitors may have promise in the treatment of diseases related to hormone receptor activation. Yet, as interest increases in the therapeutic potential of allosteric drugs, structural biologists are re-evaluating what allostery really means.

In the classic view of allostery, binding of one ligand at one site provokes a large conformational change that alters the affinity of another site for its ligand. Without denying that this conception of remote regulation has proven phenomenally successful, we can still ask whether models of this kind cover the full breadth of possibilities. Chung-Jung Tsai, Antonio del Sol, and Ruth Nussinov ask precisely that in their article now in press at the Journal of Molecular Biology (1). They find, based on an allosteric protein "benchmark", that significant backbone deformations are not an essential characteristic of allosteric effects. They therefore state that allostery might arise not only from large conformational changes, but also from changes in dynamics.

This is not a new concept—the possibility that allostery could be driven by entropic effects was articulated as early as 1984 by Cooper and Dryden (2). It seems like an odd idea, in part (I believe) because most of the analogies we use to describe allostery involve obvious changes of shape. But even stably folded proteins undergo significant fluctuations—at a fundamental level, especially when it comes to side chains, their shape is fluid. This imparts a substantial configurational entropy, which could be important in regulating binding.

Almost any binding event, irrespective of the structural dynamics of the protein or ligand, results in a decrease in the entropy of the system because the translational and rotational degrees of freedom of the protein and ligand are no longer independent. It is normal, although not always the case, that binding a ligand also significantly reduces the configurational entropy of the protein. These entropic costs are offset by energetic benefits of binding.


So, let us consider a protein that binds two ligands at different sites (see my figure cave drawing at right; CE = configurational entropy), such that neither binding event significantly alters the overall conformation. Binding of ligand A, however, might still be expected to decrease the fluctuations in the immediate vicinity of the binding site. If that's all that happens, then binding of ligand A does not alter the binding of ligand B (case I). Let us imagine, however, that the rigidification spreads from the "A site" to the "B site" (case II). Such a rigidification might pre-organize the B site; in this case the entropic cost of binding B is reduced. Thus, the binding is enhanced. Obviously, this could work the opposite way as well. It is known that binding a ligand sometimes increases the configurational entropy of a protein. If this increased entropy is communicated to the B site, then the entropic penalty for binding B and rigidifying that site is increased (case III). So depending on the system, either positive or negative allostery is possible. Tsai et al. discuss these and other models in greater detail.

Is it in fact possible for dynamic effects to be "transmitted" through a protein? As Tsai, et al. point out, previous research indicates that it is, even in relatively small globular domains. Ernesto Fuentes demonstrated that peptide binding to a PDZ domain induced changes in side-chain dynamics on the back side of the protein (3). As I discussed previously, this communication pathway has been associated with an allosteric interaction between two PDZ domains in PTP-BL. Moreover, Andrew Lee and some weirdo demonstrated that similar effects occurred in a protein of the potato inhibitor I family that lacked any discernible allosteric behavior whatsoever (4). Beyond proving that dynamic effects are transmissible, results of this kind support the idea that allosteric potential is a property of proteins generally (5).

The Tsai et al. paper is meant to reinforce and develop the concept that allostery is not solely a feature of proteins or complexes that undergo large conformational rearrangements in response to particular ligand-binding events. Small changes in backbone conformation or altered dynamics may also be responsible for allosteric effects. This may sound significantly different from present views of allostery, but in a sense it does not alter the classic interpretation. Rather, this new view points to a deficiency in the classic understanding of "conformational change". A genuine understanding of a protein's structure is not captured by a single conformation, but rather by the structural dynamics of a conformational ensemble. In some cases, i.e. the classic examples, the structural effects of ligand binding take the form of a change in the energy distribution: a particular subset of conformations becomes lower in energy and thus becomes more populated. In others, however, binding results in an elimination of possible sub-states of a given backbone conformation. Both outcomes significantly change the conformational ensemble, and both of them can produce allosteric effects.

1. Tsai, C., del Sol, A., Nussinov, R. (2008). Allostery: Absence of a change in shape does not imply that allostery is not at play. Journal of Molecular Biology DOI: 10.1016/j.jmb.2008.02.034

2. Cooper, A., Dryden, D.T. (1984). Allostery without conformational change. European Biophysics Journal, 11(2), 103-109. DOI: 10.1007/BF00276625 OPEN ACCESS

3. Fuentes, E., Der C.J., and Lee A.L. (2004). Ligand-dependent Dynamics and Intramolecular Signaling in a PDZ Domain. Journal of Molecular Biology, 335(4), 1105-1115. DOI: 10.1016/j.jmb.2003.11.010

4. Clarkson, M., Gilmore, S., Edgell, M., Lee, A. (2006). Dynamic coupling and allosteric behavior in a nonallosteric protein. Biochemistry, 45(25), 7693-7699. DOI: 10.1021/bi060652l

5. Gunasekaran, K., Ma, B., Nussinov, R. (2004). Is allostery an intrinsic property of all dynamic proteins?. Proteins: Structure, Function, and Bioinformatics, 57(3), 433-443. DOI: 10.1002/prot.20232

Read the rest...

February 22, 2008

The Evolution of Protein Folds

ResearchBlogging.orgNow online for next week's edition of PNAS is a commentary by Alan R. Davidson (1) about a paper in this week's edition of PNAS out of Matthew Cordes' group (2). Both are worth reading because they speak to a very interesting question: where do new protein folds come from?

The Roessler et al. paper doesn't address this question directly. Their initial intention was to identify the relationship between two distantly homologous proteins: P22 Cro and λ Cro. Though they both belong to the Cro repressor superfamily, these two proteins have just 25% sequence identity and significant dissimilarities in structure, as you can see on the right in the figure I have shamelessly stolen from the paper. P22 is an all α-helical structure that appears to be exclusively monomeric, while λ is an α/β structure that forms a dimer with nanomolar affinity. In an effort to bridge the structural gap between these proteins, Roessler et al. looked at group of proteins related by transitive homology. The idea, as they put it, is that
In this approach, two dissimilar sequences, A and C, are indirectly linked if a third "intermediate" sequence B exists with sufficient similarity to both A and C to imply homology with both proteins. The relationships between A and B and between B and C combine to support distant common ancestry between A and C.
Thus, they identify 3 intermediates, each with about 40% identity to its nearest neighbors, that bridge the sequence gap. Then they ask whether these sequence intermediates are also structural intermediates.

The answer is "yes", and in a somewhat surprising way. It is not the case that each step along the transitive pathway slightly increases β-strand content. Rather, while Xfaso 1 has an all-α structure very similar to P22 and is a monomer, Pfl 6—which is 40% identical (!)—has an α/β structure similar to that of λ Cro and dimerizes with ~1 mM affinity. This dissimilarity allows the authors to present some interesting ideas about the evolution of the Cro family which are summarized in their Figure 4. But what seizes Davidson's imagination is the conjunction of fairly high sequence similarity with structural dissimilarity. What makes this conjunction even more impressive is that the sequence identity is evenly distributed while the structural differences are not. The N-termini of these proteins contain structurally similar helix-turn-helix motifs, so they primarily differ in the structures of the C-termini. Yet amino-acid identity holds up across essentially the whole sequence.

Why is this such a surprise? Well, there are a variety of reasons, which Davidson outlines pretty well. It boils down to this—for a given sequence, it is generally possible to mutate a significant percentage of the residues without disrupting the fold. That is, the sequence overdetermines the structure. Consequently, proteins that have homologous or significantly identical sequences (and 40% identity would probably fall in this range) are expected to possess very similar structures. This poses a problem for protein evolution because it is expected that the initial pool of folds was rather small. If protein folds are highly resistant to disruption or alteration by mutations, it's difficult to imagine how the present enormous diversity of folds arose.

This impression is actually somewhat mistaken. It's typical to perform X→Ala mutations in these studies, and while this can occasionally produce significant cavities in a structure, it probably significantly underestimates the potential effects of a mutation at any given spot. For buried residues, size increases and the introduction of unbalanced charges (X→Trp, Asp, Lys, etc.) are mutations likely to drive the formation of new structure. For solvent-exposed residues, the introduction of bulky nonpolar side chains (X→Phe, Leu, Ile, etc.) would also be more likely to result in a novel fold than the typical approach. I have a feeling that these kinds of mutations are significant in this context, but I cannot check this because both 3bd1 and 2pij are still on hold and cannot yet be retrieved from the PDB. I may elaborate on this point when the coordinates are released to the public. For the time being I should point out that though the quantity of identity is similar between the two termini is similar, the quality is not: identical residues in the N-terminus almost all appear together, while in the C-terminus they are spread out. However, it is worth noting that what groupings of identical residues can be found in the C-terminal region tend to occur within the structural features that changed.

Davidson bears out this point when he refers to some experiments that have shown that a few mutations in key spots could change a protein's fold significantly. He seems to be unaware, however, of natural instances in which highly similar sequences produce dissimilar folds. As I have mentioned before, the upper limit on sequence identity producing dissimilar folds is known. Structural studies that Brian Volkman's group published in 2002 (3) demonstrated that the maximum sequence identity that allows for the adoption of a completely different structure is 100%. That is, given reasonable changes in solution conditions a single peptide sequence can produce two entirely different folds.

The last time I blogged on lymphotactin I discussed the implications of Brian's findings for the protein folding and protein structure prediction crowds. In the context of sequence similarity, however, the lymphotactin story also has implications for evolution as well. To a certain extent it suggests that we have been somewhat blinded by Anfinsen's dogma, in particular the assumption of the unchallenged minimum. The lymphotactin result indicates that context can be extremely important—a sequence that stably folds into one structure in one set of conditions will not necessarily maintain that structure under different conditions. In an elementary sense, we know this already, since we are aware that high temperatures and high concentrations of cosolutes such as guanidine and urea tend to unfold proteins. The important idea is that conversions between Anfinsen-like folds (i.e. folds that conditionally dominate the energy landscape) can occur within the range of conditions that can be achieved physiologically. Because the relationship between sequence and structure is not truly one-to-one, fold diversity may be much easier to achieve than we have suspected on the basis of existing structural studies.

In the end, the results of Roessler et al. provide a powerful counterpoint to conventional expectations about the relationship between sequence identity and structural homology. It appears that Cro proteins group into two kinds of structures, underscoring the well-known stability of protein folds to mutation. However, the structural discontinuity, not hinted at by sequence comparison alone, reinforces the point that fold diversity may be significantly easier to achieve than alanine-scanning mutagenesis experiments have led us to believe. Davidson nonetheless still has a significant point. In this case, it appears that transit to the new structure was relatively short, and Cro is functionally a dimer in both forms—Xfaso 1 Cro forms a dimer in the crystal, and all Cro repressors are expected to dimerize on DNA. But what happens in the transit to a completely novel structure? Can function be maintained as the protein navigates the molten-globule strewn sequence space between stable folds, and if so, how? These are questions that we will have to answer as we develop a greater understanding of the evolutionary history of biomolecules and de novo protein design.

1. Davidson, A.R. (2008). A folding space odyssey. Proceedings of the National Academy of Sciences, 105(8), 2759-2760. DOI: 10.1073/pnas.0800030105
2. Roessler, C.G., Hall, B.M., Anderson, W.J., Ingram, W.M., Roberts, S.A., Montfort, W.R., Cordes, M.H. (2008). Transitive homology-guided structural studies lead to discovery of Cro proteins with 40% sequence identity but different folds. Proceedings of the National Academy of Sciences, 105(7), 2343-2348. DOI: 10.1073/pnas.0711589105
3. Kuloglu, E.S. (2002). Structural Rearrangement of Human Lymphotactin, a C Chemokine, under Physiological Solution Conditions. Journal of Biological Chemistry, 277(20), 17863-17870. DOI: 10.1074/jbc.M200402200 OPEN ACCESS

Read the rest...

February 4, 2008

Dynamics and tunneling in soybean lipoxygenase

Blogging on Peer-Reviewed ResearchBecause even the simplest case of enzymatic catalysis involves multiple steps (i.e. association, chemistry, and dissociation), semantic issues have obscured the question of what role protein structural dynamics play. For instance, the opening of the lids in Adenylate kinase appears to be rate-limiting, but there is as yet no evidence that dynamics play any role in the phosphotranfer. Thus, one can argue that dynamics are critical (because they are rate-limiting) or irrelevant (because they contribute nothing to the chemical step). Some authors, Arieh Warshel for example, have argued that the latter finding is general—that the highly successful transition-state stabilization model leaves no place for dynamics in chemical catalysis. Some researchers, however, believe that dynamics may play a significant role in chemistry when catalysis proceeds by other routes, especially hydrogen tunneling. In this vein, Matthew Meyer, Diana Tomchick, and Judith Klinman presented evidence in last week's Proceedings of the National Academy of Sciences that protein dynamics play a significant role in the chemical step of catalysis for the protein soybean lipoxygenase (SLO-1).

Hydrogen tunneling refers to the idea that a hydrogen may in some instances "tunnel through" an energy barrier, going directly from substrate to product without wasting any time or kT in actually surmounting that barrier. A reaction that primarily reflects tunneling should have three properties. The reaction rate should vary only weakly with temperature, because tunneling mostly divorces the rate from kT. Using an Arrhenius plot, this translates to a low calculated energy of activation. Replacing the reactive hydrogen with deuterium should enormously inhibit the reaction—a large kinetic isotope effect (KIE)—because the doubling of mass makes tunneling less likely. Nonetheless, the difference in the energies of activation (ΔEa) for hydrogen and deuterium calculated from an Arrhenius plot should also be small, if tunneling is the main mechanism. Conceivably an enzyme could enhance the rate of such a reaction by using dynamics to promote favorable vibrational modes for the tunneling to occur, or to temporarily adopt a disfavored structure that puts reactive groups at a better distance for tunneling.

SLO-1 exhibits all three features, and so Klinman's group has used it as model system to try and understand whether and how enzymes promote tunneling reactions. In this case they have performed a series of mutations of isoleucine 553. In addition to existing data on WT and I553A, they analyzed the KIEs, crystal structures, and activation energies of I553L, I553V, and I553G mutants. They find that the protein as a whole is not much distorted by any of these mutations, nor do any of them appear to have significant effects on the binding pocket or substrate dissociation constant—the Kd for each mutant is around 3 μM, while WT is around 10 μM. Yet as the bulk of the mutated residue decreases, the ΔEa increases significantly. In addition, the kinetic isotope effects decrease much more sharply with temperature than for WT. Again, the magnitude of this increased temperature dependence appears to vary inversely with the bulk of the side chain at I553. From these features, and a decline in the magnitude of the pre-exponential factor, the authors argue that dynamics play an important role in encouraging the tunneling reaction.

Well, you know how I love long-range dynamic effects in proteins. But I don't quite buy it here. An argument based on negatives is never very satisfying in any case, and here we have a lot of questions. For one thing, to say that the structure hasn't changed significantly seems overly simplistic to me. The lowest-energy structure may not have been seriously deformed, but absent crystal packing and with a little extra kT around, the average structure might be different. Additionally, these structures were obtained in the absence of ligand. Even though the energetics of binding do not appear to differ significantly for these mutants, the structure of the enzyme or ligand may have changed in the bound state. Even if I accept that the existing structures argue for a completely identical binding site in the free state, and this could realistically be debated, that's no guarantee that the same is true of the bound state.

I'm not denying that the evidence is strongly suggestive, and getting direct data may be difficult given the size of the protein. However, almost all of these side chains are methylated. That means that perdeuteration of the protein in concert with specific side-chain labeling could be especially fruitful in directly observing the dynamics of the pocket during catalysis. It should also in principle be possible to label other side chains in the region to determine how mutations affect the whole area. It is likely that the substrate can also be labeled, possibly with a nucleus or tag that is poorly relaxed. T2 is likely to be a major challenge here, but not necessarily an insurmountable one, especially since the enzyme appears to be folded and active up to around 50 °C. The presence of the iron is certainly a complicating factor; however, the dynamics should be the same whether catalysis is occurring or not, so possibly it could be substituted with a non-paramagnetic metal. It wouldn't be the easiest project in the world, but it should be doable.

I think these results are intriguing, and I'd love to hear Warshel's take on them. Nonetheless, I feel I have to reserve judgment at least until dynamic changes in the pocket that seem to correlate with the kinetic observations are directly demonstrated by NMR or some other method.

Meyer, M.P., Tomchick, D.R., Klinman, J.P. (2008). Enzyme structure and dynamics affect hydrogen tunneling: The impact of a remote side chain (I553) in soybean lipoxygenase-1. Proceedings of the National Academy of Sciences, 105(4), 1146-1151. DOI: 10.1073/pnas.0710643105 - OPEN ACCESS ARTICLE

Read the rest...

January 25, 2008

What do we learn from the Protein Ensemble Method?

Blogging on Peer-Reviewed ResearchAnonymous left a comment on my post on Bruschweiller's work, referencing a couple of papers by Amarda Shehu, Cecilia Clementi, and Lydia Kavraki, the cites for which you can find at the bottom of this post. The most fascinating thing about these papers is the remarkable fidelity with which their Protein Ensemble Method (PEM) reproduces NMR-derived order parameters, 3-bond J couplings, and residual dipolar couplings. The authors demonstrate excellent correlations for ubiquitin, eglin c, Fyn SH3, Fnf10, and CI-2, and while all of these are relatively small proteins this is still a major accomplishment. Nonetheless, it is striking how little we learn from the exercise.

Keep in mind that one of the key goals of a structural biology research program is to get a veridical ensemble, i.e. an ensemble of structures that closely resembles those actually sampled by a protein under equilibrium conditions. We can learn important information from other kinds of ensembles, but the one that contains the information we are really after is the veridical ensemble.

The limitation here is intrinsic to the technique, so the technique bears some explanation. PEM utilizes an algorithm derived from robotics to move pieces of the protein. Initially, the approach was designed to map the ensemble of structures available to a loop, and I want to stress that with regards to that task I have no complaints. When given the task of mapping out the range of likely conformations of these regions this seems like an excellent approach, and the second figure of the 2006 paper seems to put this usage on fairly solid footing. The overall idea is that positioning the ends of a loop next to their anchor points is similar to solving a problem for getting a robotic arm with some number of degrees of freedom to adopt a particular pose. The authors' algorithm solves this inverse kinematic problem with a coarse-grained view of the backbone. At this point the backbone is frozen, the side chains are added back and their conformations are sampled randomly. The conformations thus generated are then subjected to energy refinement using a conventional force field. For a loop with no surroundings, this is all well and good.

The problem arises when the whole protein is subjected to the technique. This is done by using a rolling window of residues: the fragment is chosen, an ensemble defined for it while the rest of the protein is held rigid, and then the window moves to the next overlapping fragment. The various structures determined in this phase are all stored; the dynamic properties of a given residue are derived from a weighted average of all snapshots of all fragments that include that residue, with the exception that the first and last few residues of any fragment are out of bounds due to artificial restraints.

The ensemble of structures derived is therefore not veridical. Because of the fragment-replacement approach, only a single part of the protein is ever actually departing from the equilibrium or minimum-energy structure—it is unlikely that motions are actually distributed this way. Moreover, because the endpoints of all snapshots cannot be simultaneously resolved, it is not possible to assemble whole-protein conformational ensembles from the individual fragment ensembles. So, no individual snapshot is likely to reflect a significantly populated member of the ensemble, and also there is no way to collate the snapshots in such a way that the energetics of the real ensemble are accurately sampled. We thus end up with an ensemble of structures that does not reflect the set of structures actually sampled by the protein at equilibrium.

As a result, the structure that is produced can give us only limited information about the protein. For instance, this might be a reasonably reliable way to predict what sorts of deformations are possible or likely in a binding interaction. Also, PEM probably does a good job of reporting at least the lower limit of the range of the structural ensemble. However, because it does not allow for significant compensating deformations outside of the modeled region the conformations obtained probably do not cover the entire solution ensemble even for a particular fragment.

A clear implication of this work is the idea that the data are dominated by local fluctuations. That is, dynamics information derived from NMR relaxation experiments, quantitative J-coupling analysis, and RDCs primarily reflects short-range motions that do not involve major excursions from the overall structure. If this were not the case, it is unlikely that an intrinsically short-range method such as PEM could reproduce the data so well. This is not exactly a surprise, however, and the nature of PEM for the most part prevents us from learning how local motions in one region of the protein affect local motions in a distal region.

In a larger sense, however, this work reinforces the idea that the ideal approach to constructing a veridical ensemble will involve some combination of coarse-grained and all-atom approaches. The key problem here is not the computational method but the windowing. If the inverse kinematics approach used here can be extended to treat the whole protein—or at least multiple regions of the protein—simultaneously, then I think the situation improves. The question is whether this kind of algorithm will be any more efficient than MD if the whole system is in motion; I suspect at least some part of the computational savings (after what comes automatically with the coarse-graining during step one) arises from having rigid context for the fragment motions. However, this approach is also likely to be more amenable to parallelism than standard MD simulations, and because of the coarse-graining it has the ability to sample structures accessible on a timescale longer than MD can treat.

The authors of these studies imply that their future focus will be on extending this approach to larger structures; I would urge them instead to prioritize developing a way to employ PEM or a similar method without relying on fragment replacement.

Shehu, A., Kavraki, L.E., Clementi, C. (2006). On the Characterization of Protein Native State Ensembles. Biophysical Journal, 92(5), 1503-1511. DOI: 10.1529/biophysj.106.094409

Shehu, A., Clementi, C., Kavraki, L.E. (2006). Modeling protein conformational ensembles: From missing loops to equilibrium fluctuations. Proteins: Structure, Function, and Bioinformatics, 65(1), 164-179. DOI: 10.1002/prot.21060

Read the rest...

December 12, 2007

The Hierarchy of Motion in Adk

Blogging on Peer-Reviewed ResearchWell, the papers for which Ming produced his fantastic molecular artwork came out in Nature last week. There were two, as some of you might have noticed. Although the larger paper was probably more technically impressive, I'm going to skip over it and talk about the second Henzler-Wildman et al. In the interest of full disclosure, I will point out that although I did not work on this paper, those who did are in my lab. So, either I am a homer or this stuff is so exciting that I still want to write about it even after hearing it rehashed endlessly for more than a year. The question at hand is the nature of protein conformational changes. Specifically, the paper is meant to address whether motions on the timescale of nanoseconds to picoseconds are indicators of slower motions, and if so, what kind of indicators they are.

I don't mean to set up a false dichotomy here, but in order to set up this discussion it's most convenient to describe two extremes of an expected spectrum of behaviors. At one extreme of the spectrum we have a sort of "ball-and-string" model of protein behavior, in which slow motions generally reflect coherent motions of larger structural units with intervening regions of flexibility serving as "hinges". This would tend to produce the energy landscape shown at left, with two minima separated by a single large energy barrier, with perhaps a very low population of an intermediate state. Such a model predicts true two-state behavior in the slow time regime. On the fast timescale, dynamic behavior will be heterogeneous, with residues in the "ball" showing rigidity while residues in the "strings" are flexible.

As an alternative extreme, one could imagine a "jello" model of protein dynamics. In this model, domains or subdomains "wiggle" into new positions primarily through incoherent motions that occasionally produce a coherent shift. Thus, two well-populated minima might be separated by several intermediates with relatively low energy barriers (example energy diagram at right). If this is the case, one would expect the behavior in the slow time regime to be less clearly two-state (because the intermediates are all populated) and for the fast timescale dynamics to be fairly homogeneous and reflect significant flexibility. In the ball-and-string case the transition rate is governed by a monolithic energy barrier, while in the "jello" case the transition rate is limited by frustration.

I want to re-emphasize that this is not an either-or proposition. It is likely that both extremes are present in nature. Also, the ball-and-string extreme is really just a special case of the jello view. So what this paper cannot do is establish "how proteins move". However, it can establish how a particular protein moves and also indicate what techniques to apply to answer the same question in other cases.

The particular protein being studied in this paper is adenylate kinase, which I will refer to as Adk, a highly conserved protein that is present in nearly all forms of life. Its function is to convert one ATP and one AMP molecule into two ADP molecules and vice versa. Adk has two interesting properties. The first is that it is an equilibrium enzyme, which is to say that it catalyzes the phosphotransfer equally well in both directions. This is a useful feature to have in NMR studies. Additionally, although it is highly conserved, the slight differences in sequence between Adk from the intestinal bacterium E. coli (mesoAdk) and the hyperthermophile A. aeolicus (thermoAdk) produce enzymes that have very different thermal stabilities and activities.

Adk is a single fold with three subdomains: a "core" and two "lids", one of which typically binds AMP and the other ATP. It has previously been shown that the opening of these lids is the rate-limiting step of Adk catalysis, for both mesoAdk and thermoAdk. More recent research, published in a separate letter to Nature last week, indicates that these lids open and close even when the substrates are not present. The question then is whether the motion of the lids more closely resembles the ball-and-string model or the jello model. Henzler-Wildman et al. modeled thermoAdk and mesoAdk fast-timescale dynamics in order to resolve the question.

I have shamelessly stolen their first figure for your benefit. The order parameter S2 is shown on the structure of mesoAdk (A) and thermoAdk (B). Both datasets were taken at 20 °C, a temperature at which mesoAdk is very active and thermoAdk is not. Just to orient the non-NMR people, for a structured protein one expects most of the bond vectors to have order parameters between 0.8 and 0.9, so even the red residues in these structures are not really all that unusually flexible. Also note that the AMP lid is located at the right end of these structures (folding over at hinges 1-4) while the ATP lid is at the top (hinges 5-8). It should be immediately apparent that thermoAdk is more rigid than mesoAdk. A more subtle point is that many of the "hinges" identified from crystal structures have increased flexibility relative to their surroundings. Also, although the lids (especially the ATP lid) are more dynamic than the core, they retain order parameters characteristic of folded proteins.

This is in contrast to earlier research from the Meirovitch lab, which indicated substantial nanosecond flexibility in the ATP and AMP lids (LID and AMPbD in their description). The supplementary information for Henzler-Wildman et al. is freely available and contains a good discussion of the discrepancy. For those who find it too technical, it can be summed up as follows: the Meirovitch group improperly applied an isotropic global rotational diffusion model, and their SRLS results produce order parameters that are inconsistent with the domains being folded. Moreover, the correlation times they derive with their method appear to be too close together to be reliably distinguished with the NMR data they collected. The Henzler-Wildman results are consistent with the previous relaxation dispersion data for Adk and also with molecular dynamics simulations (contained in the paper). If Katie reads this, perhaps she can elaborate on the reasons why her data is superior. Commentary from the Meirovitch group is also welcome.

Let us will continue on with the paper. As I mentioned, there is a bit of MD in there which supports the NMR findings. What I found to be more interesting, however, was the way that the dynamics of thermoAdk reacted to temperature. Because thermoAdk melts at a very high temperature, it was possible to measure its dynamics all the way up to 80 °C, although glycerol had to be added to the solution at this temperature in order to bring the tumbling rate back down to a point where internal and global correlation times could be reliably separated. The results, which you will have to read the paper to see because I forgot to email myself the figure from work, indicate that as the temperature increases the order parameters decrease (expected) and become very similar to those of mesoAdk at 20 °C. This matches up nicely with the results of activity studies which indicate that the activity of thermoAdk at the high temperature is similar to that of mesoAdk at the low temperature. Because it is known that the lid opening is rate-limiting, and because local deformations at the hinges are required for the lids to open, it is likely that the increased fast-timescale flexibility of thermoAdk hinges at high temperature is directly related to the increased activity of the enzyme at that temperature.

The dynamic results in this paper seem to support the ball-and-string view of motions in Adk (by contrast, the Meirovitch results would suggest the jello view). Although the lids are not as rigid generally as the core, their order parameters are still within a reasonable range for a folded protein, and are for the most part higher than the order parameters of residues located at reasonably well-defined hinges. In Adk, motions on the μs-ms timescale appear to be enabled by flexible hotspots that lie between well-defined elements of tertiary structure, with the overall ensemble dominated by the two discrete endpoints. For this case and others like it, a combination of relaxation dispersion measurements and model-free analysis suffice to characterize the motions.

For cases which more closely match the jello view, many NMR approaches may prove difficult to implement. Relaxation-dispersion analysis of such motions is likely to be frustrated by endpoints that constitute a slim majority or plurality of the ensemble, high rates of interconversion between similar conformers, and possibly a low Δδ for these conformers. Fast-timescale studies will be difficult because the motions involved will be complex and may violate some model-free assumptions. In such cases, high-field measurements of R1 dispersion (yes, I said it) may prove especially valuable as the dominance of ωN in those measurements at least gives us a chance of distinguishing motions in the high-ns regime. Of course, the major difficulty for NMR will be if the jello motions result in intermediate exchange and excessive signal averaging — in these cases paramagnetic spin labeling or EPR studies may prove to be the only practical method.

Henzler-Wildman, K.A., Lei, M., Thai, V., Kerns, S.J., Karplus, M., Kern, D. "A Hierarchy of Timescales in Protein Dynamics is Linked to Enzyme Catalysis" Nature 450 (2007) p913-916.

Read the rest...