Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

December 16, 2009

A single residue dictates a fold

ResearchBlogging.orgAnfinsen's dogma — that the amino acid sequence of a protein uniquely determines its structure — naturally leads one to the idea that identity between amino acid sequences means identity between structures. This has proven to be a successful paradigm: sequence similarity reliably predicts structural and functional similarity. Evidence accruing in recent years, however, suggests that for small proteins, at least, this assumption may not be entirely safe. Adding to this view, in an article in PNAS this week (it's open access, so open it up), a research team led by Philip Bryan reports that they are able to generate significantly different folds with divergent functions from sequences that differ by a single amino acid.

Read the rest...

September 18, 2008

Where do new enzymes come from?

ResearchBlogging.orgBiochemists often rave about the great wonders of enzymes, lavishing praise on the prodigious rate enhancements they produce, and their exquisite positioning of functional groups. One can quite reasonably ask how such magnificently useful proteins came into being. One accurate answer, of course, is that after a couple hundred million years evolution can get almost anything right. Another answer is that most enzymes come from other proteins, via a process called gene duplication. The genetic changes that follow one of these duplications turn two copies of one protein into two completely different proteins with diverse activities.

Gene duplication events are infrequent errors of DNA replication or repair. Diploid eukaryotes such as ourselves carry two copies (or near-copies) of most genes as a matter of course, but gene duplications produce extra copies beyond that. In theory, the presence of these extra copies of a gene means that one of them can mutate freely, without the pressure of carrying out its normal job. When it drifts into a useful function, selective pressure is again applied, causing a refinement of the active site to maximize the efficiency of the new activity. The overall scheme looks something like this:

Duplication → Divergence → Refinement


It may seem incredible that a vast diversity of protein structures and activities can arise simply by making copies, even imperfect copies. However, certain quirks of the translation machinery mean that small changes in DNA can amount to enormous changes in a protein's topology. For instance, an insertion or deletion of a single base can cause a frameshift mutation, producing a protein that bears no resemblance to its progenitor despite having only 1 different base pair. Many DNA triplets that normally encode amino acids are only a single base-pair mutation away from becoming a stop codon, truncating a protein and likely changing its structure significantly. Similarly, stop codons can be easily eliminated, producing much larger proteins. In eukaryotes, point mutations near the borders between introns and exons can cause new regions of DNA to be translated into protein. Of course, drastic changes like these mostly just produce useless junk, but occasionally a novel fold or function arises.

More conservative alterations of a gene sequence can still produce significant changes. As I've mentioned before on this blog, some members of the Cro family of proteins have very high sequence identity and yet possess different structures. I also have not yet tired of reminding you that the chemokine lymphotactin has two different structures with a single sequence, either of which can be stabilized into an exclusive fold by a point mutation.

Additionally, research from the lab of John Orban shows that a mere 7 mutations are required to convert the engineered protein GA88 (PDB) into a completely different structure, GB88 (PDB) (1). These proteins were previously shown to have different folds and functions, but the contrast between the high resolution structures (shamelessly stolen figure on the right) is striking. Moreover, the Orban lab has refined this system so that the structural conversion can be effected with only three mutations, rather than seven. What all this research indicates is that the transitions that convert a sequence from one fold into another may be sharper than previously realized; even a relatively small number of fairly conservative mutations may be able to completely transform a protein's structure.

For all that, most new enzymes arising via gene duplication resemble their ancestors in identifiable ways. Often the two proteins perform the same chemical steps, and the novel function amounts to a different substrate specificity. This suggests the possibility of an alternate mechanism of gene duplication, in that a protein could evolve a novel specificity while retaining its original function. Diversifying its activities in this way would probably limit an enzyme's catalytic effect in both reactions, but a subsequent gene duplication event would allow each copy to refine its particular reaction. The scheme would look like this:

Diversification → Duplication → Refinement


The advantage of this model, from an adaptationist's perspective, is that it brings selective pressure to bear at every step. Once a new function has evolved in response to environmental conditions, duplicating the gene may provide an organism a concrete advantage. After duplication, the advantage of separately refining the two activities is obvious.

The two models are not as different as they might seem at first glance, because nearly every enzyme catalyzes two reactions anyway, that is, the forward and reverse reactions of an equilibrium. A "new" activity for a given enzyme can therefore result from something as simple as being targeted to a different cellular compartment or a change in specificity that involves an oppositely-oriented equilibrium.

The most obvious objection to the latter model is that during the period of gene sharing prior to duplication, neither protein function will be very efficient. As a matter of fact, the appearance of a new activity does not always impair an enzyme's ability to do its original job (and indeed can even enhance that activity). Still, because of the exquisite tuning of enzyme active sites we can expect that many modifications to this region will reduce catalytic power. That being the case, how might an organism survive or thrive during the gene-sharing period? The answer, which always seems obvious in retrospect, is to make more of the less efficient enzyme, as was demonstrated in a recent paper by Sean Yu McLoughlin and Shelley Copley (2).

McLoughlin and Copley took a strain of E. coli that lacked an enzyme, ArgC, that is critical for glucose metabolism. They treated these bacteria with a strong mutagen and then picked a colony that grew well on uncomplemented glucose. After showing that these bacteria had developed a novel activity equivalent to ArgC, they isolated the "new" enzyme and found that it was actually an existing enzyme, ProA, which performs similar chemistry. This enzyme had gained the ability to take over the tasks of the missing ArgC, enhancing the rate of that reaction 12-fold. The actual chemistry of these reactions was quite similar, but in gaining the ability to operate on ArgC's substrate, the activity of ProA towards its own substrate was reduced 2800-fold. The bacteria compensated for this by upregulating the production of the enzyme. A second mutation in the promoter region of the gene was helpful, but not necessary, in this respect.

Because enzymes are catalysts, a small increase in protein concentration can result in a significant increase in the availability of the reaction products. Biochemists often say, seeing a 3000-fold reduction in activity, that an enzyme is dead. The reality is that it's just slower, and a living thing can compensate for that in ways not available to an isolated reaction in a test tube. Organisms have shown that they have ways to survive what an enzymologist might see as fatal.

Of course, modern bacteria benefit from a number of well-tuned regulatory and feedback mechanisms that allow them to sense when particular metabolites are running low and to increase the production of proteins that can replenish them. Earlier, more primitive organisms might not have had these expedients available. Could they have survived gene sharing?

Too little is known about early life forms to answer such a question definitively. However, it is interesting to note that one method of making more protein is to make more of the gene. That is, the concentration of a deficient enzyme can be increased via gene duplication. By a fortuitous coincidence, a single mechanism could both enable an organism to tolerate reduced enzymatic efficiency and allow the evolutionary process to independently refine its activities.

It is also worth bearing in mind that just as ancient organisms did not necessarily resemble modern ones, ancient proteins might not have resembled the modern item. The exquisite positioning of functional groups that characterizes modern enzymes requires a rigid fold and contributes significantly to the rate accelerations they produce. However, substantial rate enhancements can still be achieved in the absence of a stiff native state.

One occasional result of mutations is the formation of a molten globule, a protein that lacks a stable fold but still exists in a collapsed state with something resembling a hydrophobic core. Although that doesn't sound particularly useful, many molten globules have enzymatic or other functional activities. Recent computational studies on a molten-globule mutant of Methanococcus jannaschii chorismate mutase suggest that realistically low energy barriers can be achieved by a broader array of structural states in these proteins (3).

Researchers from the lab of Arieh Warshel used a simplified model to sample the conformational space available to the molten globule enzyme (mMjCM) and a stably folded form of the enzyme (EcCM). As you might expect, the lowest-energy conformations are much more diverse for mMjCM than for EcCM. Roca et al. then computed the energy barrier for catalysis for conformations that closely resembled the ideal structure (region I), conformations which had most of the groups in the right general position but were significantly removed from the ideal (region II), and conformations that did not resemble the ideal at all (region III). For EcCM, only structures in region I had energy barriers low enough to plausibly allow catalysis. The molten globule, however, had energy barriers that would allow catalysis in region I and region II. You can see this in the figure below, which I shamelessly stole from their paper: the dotted orange line corresponds to a 16 kcal/mol energy barrier, what they felt to be the largest barrier reasonable for a catalyst. The results for mMjCM are on the left, EcCM on the right.



The upshot of this is that molten globules may be able to maintain catalytic power in the face of structural diversity that causes folded proteins to fail. While the stable fold produces greater rate enhancements (note that EcCM has lower energy barriers), the molten globule tolerates a wider array of structural conditions. Consequently, proteins of this kind may be much more amenable to the addition of new functions. So long as an appropriate orientation of functional groups is reasonably likely, a protein without a rigid conformation can still achieve impressive rate enhancements.

Conceivably, an early molten globule enzyme could have the ability to catalyze several different reactions, switching between the required conformations as needed, without a significant loss of catalytic power to any of them. Duplication of a multi-functional molten globule like this would allow each chemical function to be refined independently, with additional duplications and refinements giving rise to substrate specificity.

The different models of gene duplication each have their own explanatory advantages, and the available evidence suggests that new proteins and enzymatic activities have evolved (even within the last century) using both routes. As this is one of nature's favored methods of generating novel activities, so it is becoming ours. The artificial enzymes recently produced by David Baker's lab were designed onto an existing protein scaffold in what could be taken as a computational mimicry of the gene duplication process.

1. Y. He, Y. Chen, P. Alexander, P. N. Bryan, J. Orban (2008). NMR structures of two designed proteins with high sequence identity but different fold and function Proceedings of the National Academy of Sciences, 105 (38), 14412-14417 DOI: 10.1073/pnas.0805857105

2. S. Y. McLoughlin, S. D. Copley (2008). A compromise required by gene sharing enables survival: Implications for evolution of new enzyme activities Proceedings of the National Academy of Sciences, 105 (36), 13497-13502 DOI: 10.1073/pnas.0804804105

3. M. Roca, B. Messer, D. Hilvert, A. Warshel (2008). On the relationship between folding and chemical landscapes in enzyme catalysis Proceedings of the National Academy of Sciences, 105 (37), 13877-13882 DOI: 10.1073/pnas.0803405105

Read the rest...

August 30, 2008

An enzyme with a monkey's tail

ResearchBlogging.orgIt is rare, but not unheard of, for a human baby to be born with a tail. Atavism of this kind is generally understood to be the result of mutations in regulatory genes that cause an ancestral pattern of development to re-emerge. A physiological step backwards through the path of descent is often easy to recognize, because many of the evolutionary relationships are known. It should also be possible to identify atavistic events in particular molecules. For instance, one can imagine that a mutation to CLC-0 might result in a reversion to the ancestral transporter function. In a recent article in PLoS Biology, researchers from Florida State University and Brandeis University identify just such a relationship in the bi-functional enzyme inosine monophosphate dehydrogenase (IMPDH). PLoS Biology is an open-access journal, so open it up and follow along.

IMPDH plays a critical role in the synthesis of guanine nucleotides, an essential component of DNA. Two reactions take place in the active site — first, the inosine ring is oxidized to xanthosine, forming a covalent linkage with the enzyme, and then this bond is broken by a hydrolysis. The enzyme active site changes shape to carry out the reaction, bringing a catalytic arginine (R418) into position to activate the water for nucleophilic attack. Any time you see a complicated mechanism like this, it's natural to wonder how such a system could have evolved. Min et al. performed simulations and experiments to find out.

Using a crystal structure of IMPDH as a starting point, Min et al. performed hybrid QM/MM simulations in which the atoms taking direct part in the reaction were treated with quantum mechanics, and the rest of the protein was simulated using molecular mechanics. As one would expect given the enormous reduction in catalytic rate that occurs when R418 is mutated, the reaction proceeded through the arginine when the simulation had a neutral R418 side chain. The water is stabilized by two additional side chains from T321 and Y419, and reacts almost instantaneously, without the formation of a stable hydroxide intermediate. Although this is unusual, this prediction of the simulation is consistent with isotope effect experiments.

When the arginine was replaced by a glutamine in the simulation, the mechanism changed, naturally. Under these conditions, it was Y419 that activated the water for the hydrolysis, although the energy barrier was much higher (leading to a slower reaction). Again, the characteristics of the reaction indicated by the simulation line up pretty well with the results of biochemical experiments. Of course, Y419 enters the active site the same way R418 does, so the question of how the hydrolase activity could have evolved remains open.

Something very interesting, however, happens when the simulation is performed with R418 in a charged state. A fully protonated arginine will have a very hard time activating water for a nucleophilic attack. The simulation indicated that under these conditions, T321 performed this role, after being activated by a nearby glutamate (E431). T321 is adjacent to cysteine 319, which is essential for the oxidation reaction, and is not located on the mobile flap. If T321 really can catalyze hydrolysis, this would mean that it is possible that IMPDH possessed an (inefficient) hydrolysis activity before it evolved the mobile flap.

Because T321 only plays a signficant role in catalysis when R418 is protonated, blocking this pathway should result in decreased IMPDH activity at low pH. This is precisely what Min et al. observe in enzymatic assays (Figure 5) on a mutant in which E431 is mutated to glutamine. There is other experimental support as well: IMPDH enzymes that have been mutated at R418 usually have large isotope effects, which makes sense in light of the fact that the alternative T321 pathway involves the simultaneous transfer of two protons (rather than just one).

Things get even more interesting when IMPDH is compared to one of its cousins, GMP reductase. Although GMPR catalyzes a very different reaction, the C319/T321/E431 triad is also present there. This, along with other data from sequence alignment, suggests that these three residues were also present in a similar configuration in the ancestor of these modern proteins. Over time, progressive optimization of the two proteins resulted in the T321 pathway being supplanted by the more effective R418 in IMPDH, while remaining essential in GMPR.

If T321 really is a remnant of an earlier water-activating pathway, why is it conserved now that IMPDH has a much more efficient catalytic residue available? T321 is probably preserved because it stabilizes the water while it is being activated by R418. However, the other essential residue of that activating pathway (E431) is usually an inactive glutamine in eukaryotic forms of IMPDH (and some prokaryotes, as well). In these species the T321 activation pathway has been completely supplanted by the arginine pathway. Yet in the other forms of IMPDH this alternative mechanism still lingers, perhaps because of the additional activity it affords at low pH, or because it confers resistance to a particular inhibitor of the enzyme. In that sense, IMPDH's "tail" might provide an adaptive advantage quite different from that which gave rise to hydrolytic activity in the first place.

Donghong Min, Helen R. Josephine, Hongzhi Li, Clemens Lakner, Iain S. MacPherson, Gavin J. P. Naylor, David Swofford, Lizbeth Hedstrom, Wei Yang, Daniel Herschlag (2008). An Enzymatic Atavist Revealed in Dual Pathways for Water Activation PLoS Biology, 6 (8) DOI: 10.1371/journal.pbio.0060206 OPEN ACCESS
Disclaimer: Although I have little contact with Dr. Hedstrom's group, I am also working at Brandeis.

Read the rest...

June 17, 2008

Determinants and evolutionary mechanisms of homosexuality

ResearchBlogging.orgDebates over the rights of homosexuals in the United States, particularly the right to marry, often get hung up on a thoroughly inane point: whether homosexuality is "chosen" or "innate". While this may seem to be a question of moral import, it is not, and moreover it presents a false dichotomy. Like nearly all human behaviors, sexuality is too complex to be reduced to a choice or a destiny; it is neither, or both, depending on your view. However, the degree to which different factors contribute to sexuality, and the mechanisms by which they do this, are fit subjects for scientific inquiry. Two articles this week present interesting findings, sure to be distorted by all sides of the argument, that may prove enlightening in this regard. I will endeavor, along with others, to be a resource providing an unbiased view.

First, however, a plea for sanity. If science finds, by some transcension of nature, that sexual orientation is entirely chosen, or entirely innate, it does not matter to any debate over the rights of homosexuals. Men may have an innate tendency to try to spread their genes as widely as possible, but we would not forgive adultery on this basis. Toddlers have an innate tendency to become frustrated and throw tantrums, but we still make them sit in the corner. That a behavior is innate is not a basis for withholding moral judgment. And if sexual orientation is a choice? Well, we frequently forbid discrimination on the basis of chosen behaviors—religion, for instance, or political affiliation. What is truly at issue is not choice, but whether it is just to deny rights and protections to one group of citizens for no reason beyond the moral opprobrium of another group.

Thus, the question of rights for homosexuals does not depend, one way or the other, on whether people choose to be gay. I firmly believe one side of this question to be in the right, but this opinion is not informed by my scientific knowledge, because it cannot be. I would urge my readers (all three of you) to view these results strictly as what they are: interesting scientific findings related to a political question that do not support one side of the argument or the other. I would ask advocates for both sides to refrain (for once) from distorting the conclusions of these reports, not only because of the raw immorality of lying, but because by misrepresenting these findings they will have sacrificed their integrity for no gain in the debate.

There, I feel better now. On to the science!

In the first study, a team analyzed the results of a survey of Swedish twins in hopes of parsing out the relative contributions of heredity, shared environment, and unshared environment in shaping sexuality (1). Although the survey was answered by a fairly large number of twins, the authors draw their conclusions using two questions that do not directly ask for sexual orientation. The survey only requested information about actual sexual partners, and did not address homosexual feelings that might not have been acted upon. After excluding twin pairs that were opposite-sex or unclear with respect to zygosity, they had 3826 pairs to work with, of which 5% of men and 8% of women reported at least one same-sex sexual encounter. Because it is suspected that the factors influencing homosexuality may differ between the sexes, males and females were treated separately. By comparing the concordance and discordance of sexual behaviors between monozygotic and dizygotic twins it should be possible to parse out the degree to which genetics and the environment contribute to sexuality.

Despite the limited materials, the authors were able to reach some conclusions, with the caveat that the 95% confidence intervals were quite wide. For instance, for males they found that genetic factors explained 39% of the observations with respect to whether a twin had any same-sex partner in his life. However, the 95% CI on this prediction was 0% to 59%. For men, shared environmental factors appeared to contribute nothing, while unique environmental factors explained 61%. For women, it was determined that genetic factors contributed 19%, shared environment 17%, and unique environment 64%. Similar distributions were seen for comparisons of total numbers of same-sex partners. While the confidence intervals for all factors are quite large, the numbers largely agree with a previous study on Australian twins (less so with a study on American twins).

Obviously, the small sample size and broad confidence intervals on these results suggest that they should be interpreted cautiously. It should also be noted that "unique environmental factors" may run the gamut from hormone exposure in utero to childhood illness to personal experiences. Many unique environmental factors, even for twins, are just as involuntary as genetics, but some result from conscious choices of the individual (which is different from choosing to be gay). Despite their limitations, these results generally support the idea that sexual orientation results from a confluence of genetic and environmental factors.

That genetics play a role in homosexuality may seem curious, because in terms of the classic expression, "survival of the fittest", homosexuality would appear to be a non-starter. After all, a reluctance or outright inability to mate with the opposite sex would seem to result in a substantial reduction in reproductive fitness. However, contrary to what a certain ignoramus would have you believe, the Theory of Evolution has advanced substantially since the days of Darwin, and we are aware of numerous additional evolutionary mechanisms that operate alongside the law of natural selection. In the case of male homosexuality, a new paper by Camperio Ciani et al. argues that sexually antagonistic selection may be at work (2). PLoS ONE is open access, so feel free to open up the article in another window and skim it yourself.

Camperio Ciani et al. begin with the observations that male homosexuality has a matrilineal association, and that the mothers (and maternal aunts) of homosexuals are somewhat more fecund than the population at large. From these pieces of data, and from the fact that homosexuality appears to have been present at low levels in every society that has left written records, the researchers created a set of requirements for some evolutionary simulations, based on different supposed properties of the genetic factors influencing male homosexuality (GFMH). Most of the simulations failed to satisfy the parameters. In many cases (especially with single-locus traits) the GFMH either became extinct or gained too high of a frequency; in others the matrilineal association was not preserved.

Ultimately, the researchers found that the model that best fit the parameters featured two alleles (one of them X-linked), and was sexually antagonistic. What this means is that the trait increases the reproductive fitness of one sex while decreasing that of the other. For instance, a heightened sexual response to men could make women more likely to pass on their genes, while making men possessing the trait less likely to do so. Provided that the effects of this trait are balanced with respect to the population proportion of each gender, it should be possible for it to survive in a population at a relatively constant level.

This result is interesting, and provides some hypotheses that can be tested with genetics. However, it does not prove that homosexuality is genetic, or even that it has a genetic component. Like all simulations, these results merely inform us that a particular possibility is consistent with what we already know. In this case, we now know that the observed aspects of homosexuality are consistent with a 2-locus trait that is sexually antagonistic. However, this model was arrived at simply through process of elimination, and there may be some superior model or more-accurate mechanism that simply hasn't yet been tested. There is always a model we haven't thought of; sometimes that model is the right one. Moreover, as Långström et al. note, some of the data used to determine criteria for successful simulations remain controversial. Camperio Ciani et al. convincingly show how the preservation of homosexuality through evolution could happen, but that is not the same as demonstrating how it did happen. That will require a positive identification of the actual GFMH.

The results of Långström et al. indicate that any GFMH eventually identified, whether or not they materially resemble the predictions of Camperio Ciani et al., will only give rise to a heightened propensity for homosexuality. Environmental factors play a significant, perhaps even dominant, role in determining sexual orientation. Whether genetic or environmental, most factors contributing to homosexuality are involuntary, but some are chosen. If that answer doesn't satisfy you, perhaps you were asking the wrong question.

1. Långström, N., Rahman, Q., Carlström, E., Lichtenstein, P. (2008). Genetic and Environmental Effects on Same-sex Sexual Behavior: A Population Study of Twins in Sweden. Archives of Sexual Behavior DOI: 10.1007/s10508-008-9386-1

2. Camperio Ciani, A., Cermelli, P., Zanzotto, G., Brooks, R. (2008). Sexually Antagonistic Selection in Human Male Homosexuality. PLoS ONE, 3(6), e2282. DOI: 10.1371/journal.pone.0002282 OPEN ACCESS

Read the rest...

June 6, 2008

E. Coli with a more diverse palate

ResearchBlogging.orgToday wasn't a good day in the lab for me. There weren't any explosions or failed cultures, just an impossible NOESY spectrum. It's tough to determine your proline isomer when your spectrum doesn't have the characteristic peaks for the cis or trans state (note: there are no other states). At times like this, it helps to be reminded of the value of persistence, which brings me to today's paper, involving an experiment that's been going on for 20 years.

A few weeks back I mentioned Carl Zimmer's excellent book Microcosm: E. coli and the new science of life, which details the study of one of the world's best-understood organisms and the insights it has given us. E. coli has loads of useful properties that make it a great research tool, and one of these is its doubling time, which in the lab typically ranges from a few hours to as few as 20 minutes depending on conditions. This has obvious benefits for researchers who just want to use the bacteria as a means of producing protein or DNA, but it also means that scientists interested in evolution can realistically expect to use E. coli to answer questions that would require tracking a population over tens of thousands of generations. In a recent paper in PNAS, Richard Lenski's long-term evolution experiment (LTEE) has accomplished just that.

One can, for instance, ask whether evolution is random or deterministic. That is, given some specific context, is a particular outcome (or kind of outcome) inevitable, or will the evolutionary history of an organism preclude some possibilities and make others more likely? Stephen Jay Gould maintained the latter—that preceding states would provide such a significant part of the context for future mutations that rewinding time to the beginning and doing the whole thing over again might lead to a completely different world of life. This position has intuitive appeal, but it could also be the case that natural selection is so powerful that certain maximally-adapted states would be achieved regardless of evolutionary history. Obviously, we can't rewind time to test these possibilities directly, but perhaps some model system could help us.

Enter E. coli and the LTEE, an experiment that procedurally sounds very simple. What Lenski did was to found 12 populations of E. coli from two clones. Once the cultures were started, his team took part of each culture every morning and diluted it into some fresh DM25 medium, a minimal medium containing 139 µM glucose and 1.7 mM citrate. Now, E. coli loves glucose, but one of the defining characteristics of the species is that it cannot take up citrate in an aerobic environment. There is citrate inside a bacterium (as an intermediate in a metabolic pathway), and there is citrate outside the bacterium, but what is outside cannot come in. And that is how things stayed for more than 30,000 generations. A number of differences in appearance and other properties have been noted, but for a very long time the ability to eat citrate did not evolve.

Eventually, shortly after the 33,000th generation, random mutations in one of the populations gave its bacteria the ability to transport the more-plentiful citrate across their membranes. Over a very few generations, the maximum density of the cultures exploded as the bacteria gained access to the more-plentiful food source. By itself, that's pretty cool, as it represents laboratory observation of a mutation that changes a defining characteristic of a species. But this new ability also made it possible to test the importance of the previous mutations.

You see, Lenski's lab had taken samples of their E. coli every 500 generations and frozen them in glycerol at -80 °C. This sounds harsh, but this kind of deep freeze allows them to resurrect these populations for future study. And that's just what Lenski's researchers did. They pulled the samples out of deep freeze and replayed evolution for ~3700 generations for each of them, testing to see whether the ability to transport citrate under aerobic conditions evolved again.

If pre-existing context is not particularly important, one would expect that there would be some low probability of evolving citrate transport that was equal for all previous generations. On the other hand, if Gould is right, then there would be essentially no chance of evolving citrate transport for early populations, and then after some potentiating mutation occurred, a higher chance for later populations. Blount et al. show that the latter is the case. In their replays, citrate transport never evolved in populations earlier than the 20,000th generation, and only appeared regularly in replay experiments after the 30,000th generation. This would suggest that some potentiating mutation occurred before generation 20,000 that provided a genetic context in which subsequent mutations could produce citrate transport.

Obviously, one suspect for this mutation would be a change in the DNA reproduction machinery that made subsequent mutations more likely in general. It's true that some of the other populations in the LTEE had enhanced mutation rates, but not this one. A subsequent experiment tracking mutations in another gene showed no difference in mutation rate between the ancestors and the potentiated clones. So the answer is more complex. Unfortunately, the authors do not yet know exactly what mutation occurred in this time frame to potentiate citrate transport. However, our ability to examine the genes of bacteria has increased significantly since the LTEE began. Full-genome sequencing of these bacteria (now underway) should give us some powerful insights into the evolutionary history of this population.

So, would evolution play out the same way if we rewound and started again? These results suggest that it would not, and historical contingency is likely to be the case for many systems. However, it may also be true that some particular characteristics are so constrained, or confer such an enormous selective advantage, that life will take these avenues no matter what. In the case of citrate transport, the existing genetic context appears to be very significant, but generalizing this observation to other systems may not be valid. Nonetheless, Gould's point is proved even if contingency is a property of some systems. The Lenski experiment shows him to be in the right, even if it took 20 years to do it.

Carl Zimmer has a post on this article as well, complete with a few intelligent questions and some crazy rantings.

1. Blount, Z.D., Borland, C.Z., Lenski, R.E. (2008). Inaugural Article: Historical contingency and the evolution of a key innovation in an experimental population of Escherichia coli. Proceedings of the National Academy of Sciences 105(23) p. 7899-7906 DOI: 10.1073/pnas.0803151105

Read the rest...

April 18, 2008

Rapid evolution of lizards in the Adriatic

ResearchBlogging.orgEnough about Ben Stein and his lies about evolution. Let's talk some truth about evolution, namely scientific truth. The fine folks at Zooillogix brought my attention to a paper that flew under my radar at a time I may (perhaps) have been more focused on basketball. In it, researchers from Harvard and Amherst studied a population of lizards on a very small islet in the Mediterranean. In less than 40 years, this population of lizards has evolved to have significantly different morphology from its parent population, a new set of endosymbionts, and novel anatomical features rarely in related lizard species. It's an interesting case of evolution in action.

Back in 1971 10 lizards of the species Podarcis sicula were transplanted from one small islet in the South Adriatic to a nearby, somewhat smaller hunk of rock. Over a three year period from 2004-2006, Herrel et al. returned to these small islands to see what became of the lizards (1). What they found is that the transplanted lizards had taken over the second islet. The lizards from the second island were still genetically very similar to those on the originating island (see their supplementary figure 5). However, there were pronounced differences in the diet. Whereas the population on the island of origin ate very little plant matter (<10% of the total diet), the transplanted lizards appeared to subsist mostly on plant matter, with some seasonal variations up and down from 50%.

The change in diet appears to have prompted some substantial changes in morphology as well. The size and mass of the lizards on the second island is significantly greater, and the head dimensions are altered. For the smaller female lizards, this translates directly into an increase in bite force. For the male lizards, however, the differences in head size alone are not sufficient to account for the observed increase in bite force, implying that there are additional adaptations of some kind. Herrel et al. speculate that the increased bite force of the lizards helps them to eat and digest leaves. This is reinforced by the observation that structural features related to the opening of the jaw are largely unchanged.

There are additional adaptations internally. For instance, the lizards on the second island have a structure called a cecal valve (and additional anatomical changes to the cecum) that are believed to aid in the digestion of plant matter such as leaves. This is particularly interesting because this valve structure does not appear in the originating population, and is rare among related species of lizards (the suborder scleroglossa). Moreover, the hindgut of these lizards contained nematodes that are absent from the parent population, suggesting the development of a novel symbiosis. The authors note that the new morphological characteristics are present in juveniles as well as adults, suggesting that the changes involved are genetic, though further experiment is required. The authors also note some interesting changes in population dynamics and behavior that appear to have resulted from the altered eating habits of these lizards.

That these new features appeared within less than 40 years is especially striking. In less than a human lifetime this population of lizards evolved adaptations such as altered jaw morphology, as well as an apparently novel internal feature. While small populations and constrained locations can accelerate the process of evolution, it is still instructive to consider this result when discussing the "likelihood" of evolution. Significant morphological adaptations can evolve very rapidly. Imagine what the power of evolution could do given hundreds of millions of years in which to work.

Oh, wait... you don't have to.

1. Herrel, A., Huyghe, K., Vanhooydonck, B., Backeljau, T., Breugelmans, K., Grbac, I., Van Damme, R., Irschick, D.J. (2008). Rapid large-scale evolutionary divergence in morphology and performance associated with exploitation of a different dietary resource. Proceedings of the National Academy of Sciences, 105(12), 4792-4795. DOI: 10.1073/pnas.0711998105

Read the rest...

March 24, 2008

Enzymes almost as good as Ma Nature used to make

ResearchBlogging.orgBiological systems have the interesting property that most of the reactions enabling life processes are, when left to their own devices, exceedingly slow. To reach the timescales that we associate with "living", these reactions must be sped up, which requires the presence of enzymes. Because they significantly enhance reaction rates under conditions that can be encountered almost anywhere, the design of artificial enzymes is an active area of research. In two papers this month, David Baker's lab describes notable success in designing enzymes in silico to have specific activities with significant (106-fold) rate enhancement.

As a graduate student at UNC, I was fortunate to interact frequently with Richard Wolfenden, who did a great deal of work to find out just how good enzymes are at what they do (1). The fact is that there is a wide range of activities and rate enhancements. The proline isomerase cyclophilin, for instance, achieves a modest 106-fold rate increase, depending on the substrate. In contrast, arginine decarboxylase achieves an amazing rate enhancement of about 1019. Many reactions we think nothing of, such as hydrolysis of a phosphodiester bond (found in nucleic acids) would take millions of years in pure neutral water at 25° C. Of course, deviations from neutral pH and the presence of other molecules greatly enhance these rates, and obviously the same is true of changes in temperature, but this is a useful starting point for comparing enzymes to basal rates.

In the works at hand, collaborative teams involving several labs coordinated by David Baker designed enzymes to perform a novel retro-aldol reaction (2) and the Kemp elimination from 5-nitro-benzisoxazole (3) (a proton abstraction causing a ring to open). The retro-aldol paper is fascinating, particularly because of the multi-step nature of the reaction, but I'm going to focus on the Nature paper because its results are more complete, in that they implemented an appropriate wet-lab extension to the computational procedure.

The fundamental strategy of both papers is the same. For most enzymes it is believed that catalysis occurs because the transition state, the moment when the chemical reaction has the highest energy, is stabilized by the functional groups of the enzyme (see (1), among others). Using their knowledge of chemistry, the researchers of these groups predicted a transition state, and then positioned functional groups of side chains in such a way that they would stabilize this predicted state. They also placed potential bases in an appropriate geometry to attack protons as necessary. This done, they used a program based on Baker's ROSETTA to predict sequences that would fold to produce this geometry. This required a somewhat more complicated process in the case of the retro-aldol reaction due to its multiple steps.

One interesting outcome was that TIM barrels were a popular choice of this algorithm in both papers. The final results in the Röthlisberger paper are all based on backbones identified by CATH as TIM-barrel folds (explore these scaffolds at the PDB: 1thf, 1a53, 1h61, 1jcl). As the authors note, the TIM barrel is a very common catalytic scaffold in nature, in part because the central β-strands provide a convenient way to orient side-chains towards the catalytic pocket. In both papers, the structures predicted using the ROSETTA algorithm were shown to be very close to the actual result, although they only checked successful catalysts. A comparison of the failed designs to their predicted structures may be of great use in refining the computational approach.

As the above paragraph implies, the groups did in fact succeed in designing enzymes that achieved significant rate enhancements. In the case of the Kemp elimination, the eight enzymes reported had ~5x103 - 2x105 -fold increases in rate over the spontaneous reaction in a very slightly basic solution. This amounted to actual kcat (reaction rate) values of 0.006 - 0.29 s-1, which is significantly slower than is common for enzymes.

In order to improve these results, Röthlisberger et al. turned to the process that produced our own prodigious enzymes in the first place, i.e. evolution. Using a relatively standard in vitro evolution approach, they altered one of the early successes, KE07, which had a kcat of 0.018 s-1. Keep in mind, this was not the best computational design result, just one of the first that worked. This in vitro evolution procedure, in just a few rounds, produced an enzyme with a kcat of 1.37 s-1. While this is still slow for an enzyme, it represents a rate enhancement of ~1x106 over the spontaneous reaction in solution, an acceleration comparable to that of a modest enzyme like cyclophilin.

This is nowhere near a complete journey. I've already mentioned that the enzymes produced in these experiments are still quite slow in comparison to the genuine article, and the rate enhancements are still modest. The specificity of the enzymes also has yet to be proven—can these proteins distinguish their targets from a sea of similar molecules, or are they promiscuous catalysts? A further dissection of the failed designs is essential to refining the computational approach employed. More careful consideration of effects beyond the secondary shell, and (as the authors note) backbone dynamics and loop positioning may prove particularly helpful in future iterations.

So, we are not all the way to the creation of a truly proficient man-made enzyme, but this is a tremendous step in that direction. The combination of wet lab and computational approaches proved to be very successful in this case. In principle, it should be possible to incorporate all that was learned in the in vitro evolution experiments into the design algorithm from the start. We will not be designing custom catalysts for biofuel production and bioremediation tomorrow or next week. These results, however, demonstrate substantial promise for the future.

In particular, the retro-aldol paper suggests that this approach will work for multi-step reactions. However, provided that the intermediates are stable and soluble this will not be strictly necessary. So long as efficient catalysts can be designed for each step, the ability of ROSETTA to design protein-protein interfaces will make it possible to assemble functional synthetic or catabolic enzyme cassettes to achieve very complex chemistry with tremendous accelerations over basal rates.

1. Wolfenden, R., Snider, M. (2001). The Depth of Chemical Time and the Power of Enzymes as Catalysts. Accounts of Chemical Research, 34 (12), 938-945. DOI: 10.1021/ar000058i

2. Jiang, L., Althoff, E.A., Clemente, F.R., Doyle, L., Rothlisberger, D., Zanghellini, A., Gallaher, J.L., Betker, J.L., Tanaka, F., Barbas, C.F., Hilvert, D., Houk, K.N., Stoddard, B.L., Baker, D. (2008). De Novo Computational Design of Retro-Aldol Enzymes. Science, 319(5868), 1387-1391. DOI: 10.1126/science.1152692

3. Röthlisberger, D., Khersonsky, O., Wollacott, A.M., Jiang, L., DeChancie, J., Betker, J., Gallaher, J.L., Althoff, E.A., Zanghellini, A., Dym, O., Albeck, S., Houk, K.N., Tawfik, D.S., Baker, D. (2008). Kemp elimination catalysts by computational enzyme design. Nature DOI: 10.1038/nature06879

Read the rest...

March 8, 2008

Why are Xfaso and Pfl Cro so different?

A few weeks ago, when I posted on the transitive homology studies performed by the Cordes group, I promised a closer look at the structures when they became available. If you'll recall, one of the central findings of the Roessler et al. paper (1) was that the Xfaso and Pfl 6 Cro proteins, though they had 40% sequence identity, as well as an identical function, had very different structures and dimerization characteristics. The Pfl6 Cro structure is now available in the PDB, and Dr. Cordes was kind enough to send me the Xfaso Cro structure, which has been held up by some technicalities. I made the overlay of the structures to the left, with Xfaso in dark green and Pfl 6 in crimson. As you can see, the N-terminal helix-turn-helix motifs of the two molecules overlay very precisely, with some slight differences in orientation in the context helices. The C-terminal portions, of course, are completely different. How did they get to be this way?

Well, I have a few thoughts. To achieve a significant change in structure like we have here, two possibilities suggest themselves. We can destablilize one structure, or we can stabilize the other. So let's try to look at this from both angles. First, the Xfaso structure, which is on the left. The ribbon is orange for residues that don't change between the two proteins, green for mutated sites, and red over a deleted range. I've also drawn in a couple of the mutated side chains that might have an effect. For instance, at the upper right you can see a glutamate of Xfaso Cro that becomes a glycine in Pfl cro. The presence of the glycine may destabilize the helix. Lower on the helix, a solvent-exposed arginine becomes a hydrophobic leucine in Pfl; likely the structure will change to reduce the contact of the leucine with water. By the same token, a partially-buried threonine at the base of the helix gets mutated to glutamate. Not only might this put an unsolvated negative charge inside a hydrophobic region, but favorable helix-capping interactions of the threonine might be broken. Roessler et al. also point out that a pair of cysteines in the Xfaso structure are in a favorable position to form a disulfide bond; both are absent from the Pfl6 sequence.

On the other side of things, what mutations are stabilizing the new fold of Pfl 6 Cro? Check out the image below, color-coded the same way as the previous structure:
There are a couple of things going on here. First, the hydrophobic residues. The leucine that was an arginine in Xfaso Cro has indeed been buried, up at the top of the structure. Nearby, the glutamate that was a partially-buried threonine in Xfaso is now fully solvent-exposed. In the same vein, an aspartate has become an isoleucine near the middle. This new isoleucine appears to pack into the core of the companion protein near a conserved isoleucine (and a new methionine), and probably accounts to some degree for the stability of the dimer. There are a wealth of minor effects too—an alanine to arginine mutation has created a charged group that is solvent stabilized on one of the strands; down near the bottom a proline-to-arginine change probably has allowed the extension of a helix. Some other side-chains are drawn that I was taking a look at, but probably aren't critical, and of course I've probably missed a few that matter.

The R→L, T→E, and D→I mutations probably do a great deal to change the most stable conformation from all-α to α/β and strengthen dimerization in solution. However, just from looking you can never be sure what mutations are the most important. Doubtless the Cordes lab is already examining which residues are critical in altering the conformation. It will be interesting to see whether point mutations tend to move Xfaso from the all-α to the mixed structure, or whether they produce molten globules.

1. Roessler, C.G., Hall, B.M., Anderson, W.J., Ingram, W.M., Roberts, S.A., Montfort, W.R., Cordes, M.H. (2008). Transitive homology-guided structural studies lead to discovery of Cro proteins with 40% sequence identity but different folds. Proceedings of the National Academy of Sciences, 105(7), 2343-2348. DOI: 10.1073/pnas.0711589105

Read the rest...

February 25, 2008

Arieh Ben-Naim's Mystifying Decisions

I recently read Arieh Ben-Naim's book Entropy Demystified, a book I feel would have benefitted greatly from being more somewhat more focused on the task of explaining entropy than it actually was. While Ben-Naim, as promised, does a good job of reducing the second law of thermodynamics to "plain common sense", he makes confusing decisions to include arguments better meant for experts and diatribes that serve nobody. As a result I feel the book, though largely useful, may ill-serve the lay audience for which it is apparently intended, especially as regards the areas where non-scientists are most likely to encounter the concept of entropy.

In my view, the purpose of a popular book on a scientific subject is twofold. First, it should provide readers with enough good science to understand the subject as it appears (if it appears) in their lives. Second, it should provide the lay reader with a framework that allows him to interpret the words of experts. That is, the reader should come away with an understanding of what scientists mean by (in this case) entropy. Thus, popular books on any subject are not a reasonable platform to argue for paradigm shifts. First of all, you're not talking to your target audience (i.e. experts) when you do this. Second of all, the lay readers that agree with you will now be thinking of the subject in a very different way than most of the scientists they will encounter. So not only do you fail at what you're intending to do in terms of serving your new paradigm, you also fail at one basic task of the popular science book. The unfortunate thing with Ben-Naim's book is that it fails at the second task and ends up with mixed results on the first, for no real reason. He could have done just as good a job explaining entropy without introducing the new paradigm at all.

Entropy, despite its importance, is not a topic that lay people encounter very often. The transfer of heat from warmer to cooler bodies is a common example, as is the diffusion of dye in a liquid, or the mixing of two liquids or gases. Ben-Naim does an excellent job of explaining the role of entropy in these processes, building up from relatively simple dice games to complex discussions about temperature and gaseous mixtures. However, he does not do any job at all of explaining the role of entropy in reactions. Consequently, he fails to explain how apparent counteractions to entropy occur. A layman reading this book will understand why a gas expands, but not why it was possible to compress it in the first place. Also, Ben-Naim's approach and focus on mixtures and the like will also leave the reader utterly unprepared to encounter discussions of entropy in connection to the hydrophobic effect. Could the framework here introduced be extended to explain micelle formation and protein folding? I don't doubt that this is so, but Ben-Naim doesn't do it, to the detriment of the book and its readers.

What's Ben-Naim doing with all this space that he's not using on talking about the role of entropy in chemical reactions? Well, he's arguing for a view of entropy based on information, specifically defining entropy in terms of "missing information". Although this position has some merits, this book is an entirely inappropriate place to make the case for it. Why? The most pervasive popular misuse of the second law of thermodynamics is in creationist canards about the thermodynamic impossibility of evolution. These misinterpretations often rest on ideas related to information (as colloquially used) or distortions of information theory. Although Ben-Naim repeatedly states (correctly) that information has a precise scientific meaning, he never graces his readers with it. I fear this leaves them more vulnerable to creationist distortions than before they started reading. That Ben-Naim says little to emphasize the meaning of "closed system" or the other parts of thermodynamics exacerbates the problem. The discussion of information was not strictly necessary to his overall point—his excellent introduction of probability would have sufficed—and it is easily misread and misused. In my opinion, he should have held his tongue here and allowed his book targeted at the academic audience (A Farewell to Entropy: Statistical Thermodynamics Based on Information) to make the case.

Ben-Naim also devotes pages and pages to complaints about other authors' (Atkins particularly comes in for criticism) presentations of entropy. These sections contain a few good ideas, but on the whole read like petulant complaints about more popular writers. That these pages serve any real purpose at all is debatable; that they weaken Ben-Naim's case is certain. No matter how much he disagrees with these authors it would be better for him to disregard them and focus on constructing his own case than deconstructing their books. I agree with him that science writers are too eager to portray entropy as mysterious, but clearing up the mystery is sufficient. His detailed rebuttals of their language are unnecessary.

It's a pity, because Ben-Naim really does pack in some useful and important concepts. Strip away the unnecessary information theory, and his book does an excellent job of reducing the law of entropy to common sense. Entropy Demystified explains what entropy actually means for certain physical systems, and excels at building this picture up from very simple games that anyone can intuitively understand and model. That said, Ben-Naim does not tie entropy into real processes nearly as well as he could, and his insistence on representing entropy in a way that has not yet gained mainstream acceptance may mitigate his readers' ability to apprehend what scientists are saying about the subject. Moreover, the fact that this alternative presentation is "entropy as missing information" is more likely to perpetuate misunderstandings about entropy (especially with regard to evolution) among the lay audience than to clear those misunderstandings up. As such, I cannot recommend this book, despite its virtues.

Read the rest...

February 22, 2008

The Evolution of Protein Folds

ResearchBlogging.orgNow online for next week's edition of PNAS is a commentary by Alan R. Davidson (1) about a paper in this week's edition of PNAS out of Matthew Cordes' group (2). Both are worth reading because they speak to a very interesting question: where do new protein folds come from?

The Roessler et al. paper doesn't address this question directly. Their initial intention was to identify the relationship between two distantly homologous proteins: P22 Cro and λ Cro. Though they both belong to the Cro repressor superfamily, these two proteins have just 25% sequence identity and significant dissimilarities in structure, as you can see on the right in the figure I have shamelessly stolen from the paper. P22 is an all α-helical structure that appears to be exclusively monomeric, while λ is an α/β structure that forms a dimer with nanomolar affinity. In an effort to bridge the structural gap between these proteins, Roessler et al. looked at group of proteins related by transitive homology. The idea, as they put it, is that
In this approach, two dissimilar sequences, A and C, are indirectly linked if a third "intermediate" sequence B exists with sufficient similarity to both A and C to imply homology with both proteins. The relationships between A and B and between B and C combine to support distant common ancestry between A and C.
Thus, they identify 3 intermediates, each with about 40% identity to its nearest neighbors, that bridge the sequence gap. Then they ask whether these sequence intermediates are also structural intermediates.

The answer is "yes", and in a somewhat surprising way. It is not the case that each step along the transitive pathway slightly increases β-strand content. Rather, while Xfaso 1 has an all-α structure very similar to P22 and is a monomer, Pfl 6—which is 40% identical (!)—has an α/β structure similar to that of λ Cro and dimerizes with ~1 mM affinity. This dissimilarity allows the authors to present some interesting ideas about the evolution of the Cro family which are summarized in their Figure 4. But what seizes Davidson's imagination is the conjunction of fairly high sequence similarity with structural dissimilarity. What makes this conjunction even more impressive is that the sequence identity is evenly distributed while the structural differences are not. The N-termini of these proteins contain structurally similar helix-turn-helix motifs, so they primarily differ in the structures of the C-termini. Yet amino-acid identity holds up across essentially the whole sequence.

Why is this such a surprise? Well, there are a variety of reasons, which Davidson outlines pretty well. It boils down to this—for a given sequence, it is generally possible to mutate a significant percentage of the residues without disrupting the fold. That is, the sequence overdetermines the structure. Consequently, proteins that have homologous or significantly identical sequences (and 40% identity would probably fall in this range) are expected to possess very similar structures. This poses a problem for protein evolution because it is expected that the initial pool of folds was rather small. If protein folds are highly resistant to disruption or alteration by mutations, it's difficult to imagine how the present enormous diversity of folds arose.

This impression is actually somewhat mistaken. It's typical to perform X→Ala mutations in these studies, and while this can occasionally produce significant cavities in a structure, it probably significantly underestimates the potential effects of a mutation at any given spot. For buried residues, size increases and the introduction of unbalanced charges (X→Trp, Asp, Lys, etc.) are mutations likely to drive the formation of new structure. For solvent-exposed residues, the introduction of bulky nonpolar side chains (X→Phe, Leu, Ile, etc.) would also be more likely to result in a novel fold than the typical approach. I have a feeling that these kinds of mutations are significant in this context, but I cannot check this because both 3bd1 and 2pij are still on hold and cannot yet be retrieved from the PDB. I may elaborate on this point when the coordinates are released to the public. For the time being I should point out that though the quantity of identity is similar between the two termini is similar, the quality is not: identical residues in the N-terminus almost all appear together, while in the C-terminus they are spread out. However, it is worth noting that what groupings of identical residues can be found in the C-terminal region tend to occur within the structural features that changed.

Davidson bears out this point when he refers to some experiments that have shown that a few mutations in key spots could change a protein's fold significantly. He seems to be unaware, however, of natural instances in which highly similar sequences produce dissimilar folds. As I have mentioned before, the upper limit on sequence identity producing dissimilar folds is known. Structural studies that Brian Volkman's group published in 2002 (3) demonstrated that the maximum sequence identity that allows for the adoption of a completely different structure is 100%. That is, given reasonable changes in solution conditions a single peptide sequence can produce two entirely different folds.

The last time I blogged on lymphotactin I discussed the implications of Brian's findings for the protein folding and protein structure prediction crowds. In the context of sequence similarity, however, the lymphotactin story also has implications for evolution as well. To a certain extent it suggests that we have been somewhat blinded by Anfinsen's dogma, in particular the assumption of the unchallenged minimum. The lymphotactin result indicates that context can be extremely important—a sequence that stably folds into one structure in one set of conditions will not necessarily maintain that structure under different conditions. In an elementary sense, we know this already, since we are aware that high temperatures and high concentrations of cosolutes such as guanidine and urea tend to unfold proteins. The important idea is that conversions between Anfinsen-like folds (i.e. folds that conditionally dominate the energy landscape) can occur within the range of conditions that can be achieved physiologically. Because the relationship between sequence and structure is not truly one-to-one, fold diversity may be much easier to achieve than we have suspected on the basis of existing structural studies.

In the end, the results of Roessler et al. provide a powerful counterpoint to conventional expectations about the relationship between sequence identity and structural homology. It appears that Cro proteins group into two kinds of structures, underscoring the well-known stability of protein folds to mutation. However, the structural discontinuity, not hinted at by sequence comparison alone, reinforces the point that fold diversity may be significantly easier to achieve than alanine-scanning mutagenesis experiments have led us to believe. Davidson nonetheless still has a significant point. In this case, it appears that transit to the new structure was relatively short, and Cro is functionally a dimer in both forms—Xfaso 1 Cro forms a dimer in the crystal, and all Cro repressors are expected to dimerize on DNA. But what happens in the transit to a completely novel structure? Can function be maintained as the protein navigates the molten-globule strewn sequence space between stable folds, and if so, how? These are questions that we will have to answer as we develop a greater understanding of the evolutionary history of biomolecules and de novo protein design.

1. Davidson, A.R. (2008). A folding space odyssey. Proceedings of the National Academy of Sciences, 105(8), 2759-2760. DOI: 10.1073/pnas.0800030105
2. Roessler, C.G., Hall, B.M., Anderson, W.J., Ingram, W.M., Roberts, S.A., Montfort, W.R., Cordes, M.H. (2008). Transitive homology-guided structural studies lead to discovery of Cro proteins with 40% sequence identity but different folds. Proceedings of the National Academy of Sciences, 105(7), 2343-2348. DOI: 10.1073/pnas.0711589105
3. Kuloglu, E.S. (2002). Structural Rearrangement of Human Lymphotactin, a C Chemokine, under Physiological Solution Conditions. Journal of Biological Chemistry, 277(20), 17863-17870. DOI: 10.1074/jbc.M200402200 OPEN ACCESS

Read the rest...

February 18, 2008

What's that appendix for, anyway?

ResearchBlogging.orgWell, it keeps coming up, doesn't it? Famous cdesign proponentsist Dembski brought it up again recently in his list of ID "predictions" (click for epic fail). While his point was nicely deconstructed by Afarensis, I think it's worth examining the paper that attributed a function to the appendix. Just what did Bollinger et al. say about its function? On what basis did they draw their conclusions? And, I suppose most exasperatingly, what does the paper mean for evolution? Is the appendix vestigial or not, and if not, does that vindicate ID, evolution, or both?

The paper in question is an elaborately stated hypothesis, premised on the idea that
The occurrence of the appendix sporadically throughout phylogeny might suggest that the structure is evolutionarily derived for a specific function rather than merely a vestige of a once important digestive organ.

The emphasis in the above is mine, and included primarily to demonstrate that regardless of where the paper takes us it cannot be a successful unique prediction of ID, because the rationale behind the search for a function was evolutionary in the first place. But what exactly is it that the authors think the appendix is doing?

Well, they believe that the appendix is a site at which bowel biofilms are created, and a reservoir for commensal bacteria. What does this mean? Well, as we all ought be aware, we share our bodies with a significant number of bacteria who live on what we fail to digest and occasionally assist us by breaking down what we cannot and feeding bits of it to us. In recent years, as we have generally developed a greater appreciation for the role of the extracellular matrix in various aspects of biology, it has been demonstrated that our intestinal flora are to some extent supported by a rich layer of polysaccharides coating our intestinal surfaces. Bollinger et al. premise a role for the appendix in the support of this biofilm on the basis of three main pieces of evidence.

  1. The appendix is located at the proximal end of the colon, where the greatest density of biofilm can be found.
  2. Bowel biofilm is known to be associated with mucin and IgA, and the lymphoid tissue of the appendix could abundantly produce these molecules.
  3. The position of the appendix protects it from the fecal stream.
Given that biofilms benefit the host organism (us), the idea that the appendix promotes the formation of the biofilm and serves as a reservoir of friendly commensal bacteria in the event that the colon was flushed out in response to a pathogenic event (e.g. diarrhea) is an attractive attribution of function.

There are several reasons one might not find this convincing. The first of these is that biofilms are likely to be a ubiquitous feature of mammalian digestive systems. While this does not disprove the idea that the appendix promotes the formation of biofilms, it does suggest that the observed relationship between biofilm presence and the appendix might be coincidental. Similarly, success of bowel biofilms in such organisms would suggest that the abundance of lymphoid tissue in the appendix is not necessary to populate the film with the requisite immune molecules. Given the propensity for the appendix to suffer a blockage (with resulting appendicitis), one might also find proposition 3 to be tenuous, though I would welcome a rebuttal on this point from a professional gastroenterologist or anyone who studies peristalsis rigorously.

In general this hypothesis is difficult to test due to the fact that direct analogues to the human appendix in other organisms are rather rare. However, there are some experiments that can be performed, the most obvious of which is to take some animal that lacks an appendix (a carnivore of some kind?) and examine the distribution of biofilm in its colon. This would likely address whether point (1) represents a causative relationship or coincidence, and an examination of the molecular composition of the biofilms of these animals would give us good information about whether (2) is a valid reason to attribute this function to the appendix. A comparison of biofilm distribution and composition between human subjects who possess an appendix and those who lack one (at least 6% of the population, it would seem) would also indicate whether removal of the appendix deranges biofilm formation. Comparative studies of biofilm regeneration and intestinal recolonization by commensal species following diarrhea in normal and appendix-free individuals would also be of value in assessing the function of the appendix. So there are further experiments to be performed which can speak to this hypothesis.

But, let us assume for Mr. Dembski's sake that all of these experiments have taken place and yielded results which cast the best possible light on this hypothesis. Would it establish that the appendix was not vestigial? Well, no, not if you genuinely understand what vestigial means. Again, Afarensis makes some good points here, and there is also an interesting discussion of vestigiality on the appendix page at TalkOrigins. The fundamental point that you must appreciate is that vestigial does not mean useless. The word "vestigial" describes a particular kind of evolutionary history of a feature, specifically that a given physiological construct has lost the purpose it possessed in ancestral species. Vestigiality implies nothing about present function. Thus, even if it has a function in humans, the appendix would still be vestigial unless we could use it to ferment all those cellulose-laden tree leaves we eat. So it simply isn't true that the discovery of a function for the appendix means it's not vestigial—indeed, the presence or absence of a modern function doesn't speak to the vestigiality of the appendix at all.

Supposing that the hypothesis proves to be valid, it doesn't pose a real problem for evolution. Remember that the reason Bollinger et al. claim to have developed this hypothesis is to find an evolutionary rationale for the preservation of the appendix. If it turns out to provide a benefit, even in the form of a marginal improvement in biofilm robustness, then this might provide an explanation why the organ has not been lost completely. On the other hand, such an explanation is not strictly necessary; our anatomists may just be catching the appendix on the way out.

And actually, given the ubiquity of biofilms, validation of this hypothesis would represent a(nother) philosophical disaster for the intelligent design conjecture. If the appendix truly provides a significant benefit to biofilm formation, then it seems logical to ask why the designer included it in so few mammals, of such disparate kinds. Why not put it in all of them? If it's not important for that, then why put it in any? The apparently random distribution (and dissimilarity) of appendices in mammals is consistent neither with a purposeful insertion of elements known to be of benefit, nor with a purposeful trimming of elements known to be useless. Evolution, on the other hand, is perfectly at home with this kind of convergent funny business. In this sense, it doesn't matter whether a function is ever found for the appendix or not. Vestigiality aside, the existence of the appendix cannot be explained by the design conjecture if it is useless, nor can it be explained by design conjecture if it possesses a function, as this hypothesis suggests, that would be of benefit to all animals, including those that lack one.

Bollinger, R.R., Barbas, A., Bush, E., Lin, S., Parker, W. (2007). Biofilms in the large bowel suggest an apparent function of the human vermiform appendix. Journal of Theoretical Biology, 249(4), 826-831. DOI: 10.1016/j.jtbi.2007.08.032

Read the rest...

February 7, 2008

What was Leslie Orgel saying?

Blogging on Peer-Reviewed ResearchWell, my first impression on reading Leslie Orgel's recent (posthumous) essay in PLoS Biology was that it was just the sort of rich quote-ore that creationists and cdesign proponentsists cannot resist mining. Indeed, it was not long before famous cdesign proponentsist Casey Luskin began distorting the article in a way that served his personal beliefs. Others have dissected the failures of Mr. Luskin's analysis, and of course volumes were written about his putative misuse of the Research Blogging icon. In all the clamor condemning Mr. Luskin, Orgel's actual paper seems to have gotten lost. I thought I might write about it.

Orgel's article itself doesn't contain any new research, and those who are keeping track might notice similarities to a previous essay on the same topic from 2000. The issue at hand is the origin of life, which owing to our inability to travel back in time is one of the most difficult subjects in biochemical research. It is almost impossible to know exactly what chemicals were present in the prebiotic milieu, what minerals and surfaces were available to perform catalysis, and which of the many possible environments was the one in which life actually arose. As such, the field is highly speculative. Orgel's essay is something of a reaction to this.

It is popular to speak of an "RNA world" of primitive organisms in which most or all of the functions currently performed by proteins were instead performed by RNA. This model has the advantage that RNA can have both informational content and catalytic activity, but it demands the question of where the RNA came from. Some elements of nucleic acids can form nonenzymatically from simple chemicals likely to be present in the prebiotic milieu, but the transition from adenine and ribose to a diverse oligosaccharide with a phosphate backbone is not trivial. If, however, enormous autocatalytic cycles existed on the prebiotic Earth, that would resolve many of these objections.

In his essay, Orgel seemingly meant to quell enthusiasm (I use the term in its pejorative sense) for this idea, not because he believes it is intrinsically false, but because it doesn't really make things easier for us. He takes as an example the reverse citric acid cycle, pointing out that although it is extremely useful, and moreover is autocatalytic, it requires numerous and diverse chemical activities, many of which could go awry if the wrong substrate were used for a step. The assertion that this cycle underpins the existence of life requires appropriate colocal catalysts acting with sufficient efficiency to keep the cycle running. Moreover, it requires that unproductive (or toxic) side-reactions not drain away the reactants at any step. That is, the catalysts present must discriminate between different possible reactants, so as not to break down components inappropriately before they advance in the cycle. It is not immediately apparent that this is possible (certainly it does not appear to be probable), and Orgel was not convinced by several of the attempts to justify prebiotic catalytic cycles. This article summarized his reasons for skepticism.

This attitude has been misrepresented as meaning that Orgel believed (A) that prebiotic autocatalytic cycles were impossible, and (B) that important life cycles are irreducibly complex. Both propositions are false, and it is also false that Orgel believed them.

Part of the rebuttal to proposition A is referenced in Orgel's paper itself, when he mentions Arthur Weber's recent work in which a reaction of various trioses with ammonia gave autocatalytic products. Although the cycle at work in that instance is not yet understood, Orgel points to it as being particularly promising, in part because it requires no additional catalyst, and in part because it yields high-energy carbon compounds of a kind that might be useful substrates for life.

Granted, the reactions of the Weber experiment may not be directly analogous to anything observed in modern organisms, but I feel that this should not be seen as a problem. After all, why would a primitive organism evolve an activity to perform catalysis already occurring naturally? Rather, one would expect that primitive organisms, whatever their informational and structural characteristics, used the products of autocatalytic cycles as raw material for more exotic activities, which were later repurposed to the production of raw materials as a way to gain a competitive edge or to respond to resource scarcity. If, additionally, very complex abiotic cycles similar to the TCA cycle existed to produce more exotic or useful substrates, it ceases to be quite as problematic if they are not highly efficient or specific. Indeed, these sorts of shortcomings would likely serve as a basis for establishing selective advantage for those organisms that evolved compensating protein catalytic activities.

It is not necessary that abiotic cycles intended to serve as a metabolic origin of life work perfectly, or even that they involve products and intermediates similar to those used by contemporary organisms. They need only be efficient and specific enough for life to get started, with intermediates and products that are useful enough for primitive organisms to benefit from them. Orgel's objection is not that this latter situation is impossible; rather he feels that it is sufficiently implausible that actual evidence is needed, rather than hopefulness and modeling.

The sharp-eyed will (hopefully) note the influence of Orgel's rules in the above, especially the First Rule: "Whenever a spontaneous process is too slow or too inefficient a protein will evolve to speed it up or make it more efficient." The rebuttal to proposition B, of course, comes in the form of Orgel's Second Rule: "Evolution is smarter than you are." The belief that any biological systems are irreducibly complex results merely from a failure to understand the enormous problem-solving potential of selection combined with random mutation and millions upon millions of years of time to work. Certainly Orgel would never have accepted that evolution could not produce these complicated cycles once given a chance, or that these cycles could not be broken down in a useful way for primitive organisms. Orgel's point is not that any of this is impossible, it is that it must be shown to be plausible:
The most serious challenge to proponents of metabolic cycle theories—the problems presented by the lack of specificity of most nonenzymatic catalysts—has, in general, not been appreciated. If it has, it has been ignored. Theories of the origin of life based on metabolic cycles cannot be justified by the inadequacy of competing theories: they must stand on their own.

Orgel's paper does not rule out the possibility of autocatalytic cycles, nor does it assert or even imply that some mystical intervention is necessary to explain how life came about. Orgel's paper is a demand for evidence. It is not enough to say that autocatalytic cycles solve the problem of biogenesis, nor even that they are more plausible than other explanations. Negative arguments do not suffice; a positive argument must be made and backed up by simulations, reconstructions, and experimental evidence. Orgel points favorably to several promising avenues of research in this regard.

Intelligent design is not among those avenues, and it is certain that Orgel would reject it. Having only a (flimsy) negative case in its favor, lacking even the virtue of plausibility, and ignoring his second rule, ID must be seen as weaker than any extant metabolic or genetic theories, which at least are not based on magic. Orgel's overall point is a valuable one. It is disgusting and disgraceful that, with flagrant disregard for his thinking, Orgel's essay should be misrepresented as supporting "if pigs could fly" views of the origin of life.

Primary citation:
Orgel, L.E. (2008). The Implausibility of Metabolic Cycles on the Prebiotic Earth. PLoS Biology, 6(1), e18. DOI: 10.1371/journal.pbio.0060018 OPEN ACCESS

Other peer-reviewed articles referenced:
Orgel, L.E. (2000). Self-organizing biochemical cycles. Proceedings of the National Academy of Sciences, 97(23), 12503-12507. DOI: 10.1073/pnas.220406697 OPEN ACCESS
Roy, D., Najafian, K., von Rague Schleyer, P. (2007). Chemical evolution: The mechanism of the formation of adenine under prebiotic conditions. Proceedings of the National Academy of Sciences, 104(44), 17272-17277. DOI: 10.1073/pnas.0708434104
Weber, A.L. (2007). The Sugar Model: Autocatalytic Activity of the Triose-Ammonia Reaction. Origins of Life and Evolution of Biospheres, 37(2), 105-111. DOI: 10.1007/s11084-006-9059-9

Read the rest...