Gene duplication events are infrequent errors of DNA replication or repair. Diploid eukaryotes such as ourselves carry two copies (or near-copies) of most genes as a matter of course, but gene duplications produce extra copies beyond that. In theory, the presence of these extra copies of a gene means that one of them can mutate freely, without the pressure of carrying out its normal job. When it drifts into a useful function, selective pressure is again applied, causing a refinement of the active site to maximize the efficiency of the new activity. The overall scheme looks something like this:
Duplication → Divergence → Refinement
It may seem incredible that a vast diversity of protein structures and activities can arise simply by making copies, even imperfect copies. However, certain quirks of the translation machinery mean that small changes in DNA can amount to enormous changes in a protein's topology. For instance, an insertion or deletion of a single base can cause a frameshift mutation, producing a protein that bears no resemblance to its progenitor despite having only 1 different base pair. Many DNA triplets that normally encode amino acids are only a single base-pair mutation away from becoming a stop codon, truncating a protein and likely changing its structure significantly. Similarly, stop codons can be easily eliminated, producing much larger proteins. In eukaryotes, point mutations near the borders between introns and exons can cause new regions of DNA to be translated into protein. Of course, drastic changes like these mostly just produce useless junk, but occasionally a novel fold or function arises.
More conservative alterations of a gene sequence can still produce significant changes. As I've mentioned before on this blog, some members of the Cro family of proteins have very high sequence identity and yet possess different structures. I also have not yet tired of reminding you that the chemokine lymphotactin has two different structures with a single sequence, either of which can be stabilized into an exclusive fold by a point mutation.
Additionally, research from the lab of John Orban shows that a mere 7 mutations are required to convert the engineered protein GA88 (PDB) into a completely different structure, GB88 (PDB) (1). These proteins were previously shown to have different folds and functions, but the contrast between the high resolution structures (shamelessly stolen figure on the right) is striking. Moreover, the Orban lab has refined this system so that the structural conversion can be effected with only three mutations, rather than seven. What all this research indicates is that the transitions that convert a sequence from one fold into another may be sharper than previously realized; even a relatively small number of fairly conservative mutations may be able to completely transform a protein's structure.For all that, most new enzymes arising via gene duplication resemble their ancestors in identifiable ways. Often the two proteins perform the same chemical steps, and the novel function amounts to a different substrate specificity. This suggests the possibility of an alternate mechanism of gene duplication, in that a protein could evolve a novel specificity while retaining its original function. Diversifying its activities in this way would probably limit an enzyme's catalytic effect in both reactions, but a subsequent gene duplication event would allow each copy to refine its particular reaction. The scheme would look like this:
Diversification → Duplication → Refinement
The advantage of this model, from an adaptationist's perspective, is that it brings selective pressure to bear at every step. Once a new function has evolved in response to environmental conditions, duplicating the gene may provide an organism a concrete advantage. After duplication, the advantage of separately refining the two activities is obvious.
The two models are not as different as they might seem at first glance, because nearly every enzyme catalyzes two reactions anyway, that is, the forward and reverse reactions of an equilibrium. A "new" activity for a given enzyme can therefore result from something as simple as being targeted to a different cellular compartment or a change in specificity that involves an oppositely-oriented equilibrium.
The most obvious objection to the latter model is that during the period of gene sharing prior to duplication, neither protein function will be very efficient. As a matter of fact, the appearance of a new activity does not always impair an enzyme's ability to do its original job (and indeed can even enhance that activity). Still, because of the exquisite tuning of enzyme active sites we can expect that many modifications to this region will reduce catalytic power. That being the case, how might an organism survive or thrive during the gene-sharing period? The answer, which always seems obvious in retrospect, is to make more of the less efficient enzyme, as was demonstrated in a recent paper by Sean Yu McLoughlin and Shelley Copley (2).
McLoughlin and Copley took a strain of E. coli that lacked an enzyme, ArgC, that is critical for glucose metabolism. They treated these bacteria with a strong mutagen and then picked a colony that grew well on uncomplemented glucose. After showing that these bacteria had developed a novel activity equivalent to ArgC, they isolated the "new" enzyme and found that it was actually an existing enzyme, ProA, which performs similar chemistry. This enzyme had gained the ability to take over the tasks of the missing ArgC, enhancing the rate of that reaction 12-fold. The actual chemistry of these reactions was quite similar, but in gaining the ability to operate on ArgC's substrate, the activity of ProA towards its own substrate was reduced 2800-fold. The bacteria compensated for this by upregulating the production of the enzyme. A second mutation in the promoter region of the gene was helpful, but not necessary, in this respect.
Because enzymes are catalysts, a small increase in protein concentration can result in a significant increase in the availability of the reaction products. Biochemists often say, seeing a 3000-fold reduction in activity, that an enzyme is dead. The reality is that it's just slower, and a living thing can compensate for that in ways not available to an isolated reaction in a test tube. Organisms have shown that they have ways to survive what an enzymologist might see as fatal.
Of course, modern bacteria benefit from a number of well-tuned regulatory and feedback mechanisms that allow them to sense when particular metabolites are running low and to increase the production of proteins that can replenish them. Earlier, more primitive organisms might not have had these expedients available. Could they have survived gene sharing?
Too little is known about early life forms to answer such a question definitively. However, it is interesting to note that one method of making more protein is to make more of the gene. That is, the concentration of a deficient enzyme can be increased via gene duplication. By a fortuitous coincidence, a single mechanism could both enable an organism to tolerate reduced enzymatic efficiency and allow the evolutionary process to independently refine its activities.
It is also worth bearing in mind that just as ancient organisms did not necessarily resemble modern ones, ancient proteins might not have resembled the modern item. The exquisite positioning of functional groups that characterizes modern enzymes requires a rigid fold and contributes significantly to the rate accelerations they produce. However, substantial rate enhancements can still be achieved in the absence of a stiff native state.
One occasional result of mutations is the formation of a molten globule, a protein that lacks a stable fold but still exists in a collapsed state with something resembling a hydrophobic core. Although that doesn't sound particularly useful, many molten globules have enzymatic or other functional activities. Recent computational studies on a molten-globule mutant of Methanococcus jannaschii chorismate mutase suggest that realistically low energy barriers can be achieved by a broader array of structural states in these proteins (3).
Researchers from the lab of Arieh Warshel used a simplified model to sample the conformational space available to the molten globule enzyme (mMjCM) and a stably folded form of the enzyme (EcCM). As you might expect, the lowest-energy conformations are much more diverse for mMjCM than for EcCM. Roca et al. then computed the energy barrier for catalysis for conformations that closely resembled the ideal structure (region I), conformations which had most of the groups in the right general position but were significantly removed from the ideal (region II), and conformations that did not resemble the ideal at all (region III). For EcCM, only structures in region I had energy barriers low enough to plausibly allow catalysis. The molten globule, however, had energy barriers that would allow catalysis in region I and region II. You can see this in the figure below, which I shamelessly stole from their paper: the dotted orange line corresponds to a 16 kcal/mol energy barrier, what they felt to be the largest barrier reasonable for a catalyst. The results for mMjCM are on the left, EcCM on the right.

The upshot of this is that molten globules may be able to maintain catalytic power in the face of structural diversity that causes folded proteins to fail. While the stable fold produces greater rate enhancements (note that EcCM has lower energy barriers), the molten globule tolerates a wider array of structural conditions. Consequently, proteins of this kind may be much more amenable to the addition of new functions. So long as an appropriate orientation of functional groups is reasonably likely, a protein without a rigid conformation can still achieve impressive rate enhancements.
Conceivably, an early molten globule enzyme could have the ability to catalyze several different reactions, switching between the required conformations as needed, without a significant loss of catalytic power to any of them. Duplication of a multi-functional molten globule like this would allow each chemical function to be refined independently, with additional duplications and refinements giving rise to substrate specificity.
The different models of gene duplication each have their own explanatory advantages, and the available evidence suggests that new proteins and enzymatic activities have evolved (even within the last century) using both routes. As this is one of nature's favored methods of generating novel activities, so it is becoming ours. The artificial enzymes recently produced by David Baker's lab were designed onto an existing protein scaffold in what could be taken as a computational mimicry of the gene duplication process.
1. Y. He, Y. Chen, P. Alexander, P. N. Bryan, J. Orban (2008). NMR structures of two designed proteins with high sequence identity but different fold and function Proceedings of the National Academy of Sciences, 105 (38), 14412-14417 DOI: 10.1073/pnas.0805857105
2. S. Y. McLoughlin, S. D. Copley (2008). A compromise required by gene sharing enables survival: Implications for evolution of new enzyme activities Proceedings of the National Academy of Sciences, 105 (36), 13497-13502 DOI: 10.1073/pnas.0804804105
3. M. Roca, B. Messer, D. Hilvert, A. Warshel (2008). On the relationship between folding and chemical landscapes in enzyme catalysis Proceedings of the National Academy of Sciences, 105 (37), 13877-13882 DOI: 10.1073/pnas.0803405105


1 comment:
Thank you for this post! just what I needed!
Post a Comment