Showing posts with label selenocysteine. Show all posts
Showing posts with label selenocysteine. Show all posts

Wednesday, November 6, 2013

Rewriting the Genetic Code

Biology concepts – DNA, RNA, tRNA, nonstandard nucleotides, codon, anticodon, genetic code, selenocysteine, isodecoder, mitochondria


Just looking the Imperial Hotel in Tokyo doesn’t really give us an idea
of why they inspired Frank Lloyd Wright’s son to invent Lincoln Logs.
 It was the interlocking beams of the basement in which his vision was
born. They were supposed to protect the hotel from earthquake
damage. It worked. In 1923, the same year the hotel was finished,
there was a great earthquake in Tokyo and the Imperial was one of
the few buildings that survived. It also survived the bombings of
WWII. So they tore it down in 1968.
In 1916 John Lloyd Wright invented Lincoln Logs. The construction set was based on his memory of the Imperial Hotel in Tokyo, an edifice designed by his father, Frank Lloyd Wright. The construction set had specific pieces that fit together in a specific way.

The first edition of Lincoln Logs, sold in 1918, gave instructions for building Abraham Lincoln’s boyhood home and Uncle Tom’s cabin. The parts were commensurate for building those structures. Each set of instructions called for the small pieces to be put together in a certain order so that the resulting product conferred a meaning – this is where Lincoln grew up or this is where Tom lived.

DNA and RNA are similarly constructed. There are a few pieces (nucleotides A, C, G, T, and U) that can be used to build different structures. Each small piece can be joined with other small pieces to become part of the whole structure, a structure with meaning. In the case of DNA and mRNA, three nucleotides in a row can confer meaning for one protein building block. The entire series of nucleotides then has the meaning of an entire protein.

The three nucleotide codons relate to a certain amino acid building block to be inserted into a growing protein. This code, the genetic code, gives meaning to the string of DNA nucleotides in genes and the string of nucleotides in the mRNA transcribed from the gene. This is usually where our learning about nucleic acids ends.


The top left picture is Marshall Nirenberg, the initial decoder of the
genetic code. The right photo is Robert Holley, discoverer of tRNA.
Below is the genetic code in graphic style. The four large letters
represent the possible first bases of a codon (in mRNA). The light
yellow letters are the possible 2nd bases, and the darker yellow letters
are the final possibilities. Outside are the amino acids that are coded
by the individual codons. Note that most have more than one codon
and some have codons that begins with different letters, like serine
at 1:00 and 8:30.
The history of the genetic code is worth knowing, as is the history of about every part of science. I often use history to illustrate points in the blog. It is said that those who ignore history are doomed to repeat it, and science has its own version of this axiom, “Six months in the lab can save you a whole afternoon in the library.” Think about it. And besides saving you from repeating others' work, knowing history helps you ask better questions.

But I digress – let’s talk briefly how we decoded the pathway of gene to protein. It begins with Watson and Crick publishing the structure of DNA in 1953. We knew how the different bases could be ordered, but we still didn’t know how they called for a specific amino acid sequence.

In 1955, Francis Crick thought he had an idea about how it might occur, but he didn’t have all the players. He called his idea the Adapter Hypothesis. What he was missing was the adapter, the piece that he said carried amino acids and put them in the correct order.

One neat trick came from George Gamow, a nuclear physicist best known for his role in theorizing the Big Bang (the birth of elements from a cosmic explosion, not the TV show). We had four nucleotides to encode information and 20 (you and I know there are 22) amino acids to be coded for. He used some “way beyond me” math to determine that the most efficient mechanism would have three nucleotides code for one amino acid.

This was followed by an interesting experiment done by Marshall Nirenberg at the National Institutes of Health near Washington, DC. He made a synthetic RNA of a single nucleotide (UUUUU….). He then combined this with the innards of a bunch of cells (cell lysate) so that everything needed to make a protein would be present. He detected a peptide of phenylalanine amino acids. What is more, there were 1/3 as many amino acids as there were nucleotides!

So UUU coded for phenylalanine. This was followed by many more experiments using different sequences of nucleotides, and the code was decoded. Along with this knowledge came the discovery of tRNA by Robert Holley in 1965. This RNA combined an anticodon sequence to recognize a codon on mRNA and carried the appropriate amino acid at the other end. The tRNA was Crick’s adapter, and perhaps the code would have discovered years earlier if the adapter had been pursued in earnest.


The process of turning an mRNA into a protein involves the ribosome and
the tRNAs. When an mRNA is bound to a ribosome, the three nucleotides
in the codon (pink letters) match with three letters of a tRNA anticodon
(blue letters). Different tRNAs will drift in and out until the right one is
bound. The tRNA has the amino acid (aa) bound to the end opposite the
anticodon. If this is the first position of the peptide, it will occur in the P
(peptidyl) site. The second tRNA will be added to the A (acceptor) site and
the ribosome will shift as it creates a peptide bond between the two aa’s.
The shift puts the first aa in the E (exit) site to reelase the tRNA, the 2nd aa
goes to the P site, and the A site is open for the next tRNA.
There are 64 possible codons that can be made from four nucleotides (4 x 4 x 4), but Holley found fewer than 64 tRNAs, one for each codon. Even I know that this kind of math doesn’t work. It turned out that the genetic code was degenerate; more than one codon calls for a particular amino acid. Most amino acids have 2-4 codons assigned to them (we have talked about the exceptions to that rule).

In most cases, codons that call for the same amino acid have the same first two nucleotides; it’s the third position (wobble position) that varies. It was discovered that the third position of the anticodon binds to the DNA very loosely, so the codon/anticodon binding is usually determined by the first two nucleotides. This allows a single tRNA to recognize more than one codon.

It turns out that there are 40-55 different tRNAs, depending on the organism. Why so many? As an example, arginine is coded for by several codons (CGG, CGA, CGC, CGU, AAC, and AAU). It is impossible for one tRNA to recognize both AAC and CGG, so there must be more than one tRNA for arginine.

Serine and leucine are like this as well, and there are most certainly some amino acids whose tRNAs can’t bind to all four possible nucleotides in the wobble position (like glycine), so they would need more than one tRNA. These are the isodecoder tRNAs (different anticodons, but code for same amino acid).

There are also different isodecoder tRNA genes, having different sequences outside the anticodon, but code for the same amino acid. Humans have about 274 genes for our 55 different tRNAs. This implies that the different sequences might have some functions other than just helping to add the right amino acid to a growing peptide sequence.

A 2010 minireview talked about those possible tRNA functions. In one discussed study, a cleaved tRNA is shown to have increased expression when cells are proliferating. Reducing the levels of this cleavage product reduced the rate of cell division. In another study, a tRNA cleavage product silenced the expression of a specific gene. I’ve said it before: nature abhors a unitasker.

UGA, UAA, and UAG are the most common stop codons (see the text for the
exceptions). When the stop codon ends up in the A site, no tRNA fits properly,
but a releasing factor (RF) can be bound. There are at least two RF, RF-1
recognizes UAA and UAG; RF-2 recognizes UAA and UGA. When bound they
cause the ribosome to fall apart.

There are also three codons that don’t code for an amino acid. These are the stop codons that tell the ribosome to stop making the protein and release it.

So we have coding codons and noncoding codons. Experiments in other organisms in the 1960’s and 1970’s indicated that all life uses the same genetic code, making it the universal genetic code. And here begins the exceptions.

The genetic code is almost universal. Considering how many genes from how many organisms there are, the number of exceptions is relatively low. But they are still too numerous for us to talk about them all. That doesn’t mean we should talk about a few of the most interesting.

Mitochondria are the source of many of the exceptions. The endosymbiotic theory states that a bacterium was engulfed by an archaea and they agreed to allow each other to do what they do best. These engulfed bacteria became mitochondria and chloroplasts. But they didn’t always follow the same path.


We are finding that tRNAs can have multiple functions. 1- This is the
usual route, the tRNA codes for an amino acid in a growing peptide.
2- Some tRNAs code for the carry the same amino acid, but have
differences in structure. The change in structure means they don’t
bind the amino acid, so they are free to do other things. 3 and 4- These
non-aa bound tRNAs may be used for regulating expression of specific
genes, usually in the end of the gene. 5- there are probably functions
we don’t know yet.
Remember that mitochondria have their own genomes and machinery for transcribing DNA to mRNA, and translating mRNA to protein. This includes their own set of tRNAs. Since they are all packaged in a closed system, there is no demand for mitochondria to use the same genetic code as nuclear genes. And in many cases, they don’t.

In animal and protist mitochondria, but not plants, the stop codon UGA instead codes for the amino acid tryptophan. You’d think that this would leave them with just two possible stop codons, and some do. But in vertebrates, the codons AGA and AGG (usually code for arginine) have been converted to stop codons. So we actually have four mitochondrial stop codons.

Furthermore, animal mitochondria have switched up another codon; AUA codes for methionine instead of isoleucine. In yeast mitochondria, all the CU_ codons code for threonine instead of leucine. Again I ask… why? I ask that a lot. Not so much why the genetic code has changed in mitochondria, but why it hasn’t in plants.  You tackle that one on your own.

Nuclear genes have far fewer exceptions to the universality of the genetic code. A protist or two have converted two stop codons to code for glutamine, and the bacterium Mycobacterium capricolum has converted the stop codon UGA to a tryptophan codon. Beyond that, we have couple exceptions we have already discussed a bit, selenocysteine and pyrrolysine.

The interesting story is selenocysteine (SeC). We said that it is coded for by a stop codon plus a special stem/loop structure downstream called the SECIS structure. This makes it the 21st amino acid. If it is coded for, even indirectly, it’s going to need a tRNA. In this case, a serine tRNA is modified in a two-step process to carry a SeC.


These are two marine ciliate protist Euplotes crassus organisms
undergoing sexual reproduction, a marine ciliate protist. They are
interesting for many reasons, but one is that they use a slight
variation of the genetic code, and the other reason has to do with
something called a frameshift. The codons are read in 3’s, but some
genes in E crassus require a shift in the reading frame to produce
the correct protein. This means that they go along a 3, 3, 3, 3, then
the ribosome has to move 1 nucleotide over, and then it starts
reading 3, 3, 3, 3 again. The one nucleotide doesn’t code for
anything, but must be there to change the reading frame.
A recent paper identified that the stop codon UGA in Euplotes crassa codes for both Sec and cysteine. Which one gets put in to the growing peptide is based on how far the site is from the SECIS structure.

The same group has a new paper that says humans can also end up with cysteine in the Sec site (originally a UGA stop codon). How can these two examples of cysteine in a Sec site take place, especially since the cysteine and SeC tRNAs are completely different?!

It turns out that it's the levels of selenium and a molecule called thiosulfate (SPO4) that is important for converting other amino acids to cysteine. In some cases, the serine tRNA can be made into a cysteine tRNA instead of a SeC tRNA. So here we have a case of a UGA stop codon converted to a Sec codon then converted to a cysteine codon. Exceptional.

Next week, we can finish up nucleic acid exceptions. Do you think A, G, C, T, and U are it when describing nucleotides? Not even close.




Xu XM, Turanov AA, Carlson BA, Yoo MH, Everley RA, Nandakumar R, Sorokina I, Gygi SP, Gladyshev VN, & Hatfield DL (2010). Targeted insertion of cysteine by decoding UGA codons with mammalian selenocysteine machinery. Proceedings of the National Academy of Sciences of the United States of America, 107 (50), 21430-4 PMID: 21115847

Thoru Pederson (2010). Regulatory RNAs derived from transfer RNA? RNA DOI: 10.1261/rna.2266510



For more information or classroom activities, see:

Genetic code –

Isodecoder tRNAs –
http://ymalblog.blogspot.com/2011/10/misfolded-human-trna-isodecoder-binds.html


 

Wednesday, August 28, 2013

So Many From So Few

Biology concepts – protein, amino acids, non-standard amino acids, peptide bond


Severe dietary protein deficiency leads to distinct
symptoms, and if not resolved, death. Called kwashiorkor
(Ghanan word meaning “disease from second born”),
the deficiency leads to changes in osmotic potential in
the bodies cells as compared to their blood.
Hypoalbuminemia (low levels of the blood protein
albumin) lead to fluid leaving the vessels and accumulating
in the abdomen, called ascites. This often occurs when infants
stop nursing (like when a second child is born); they take in
enough calories but not enough protein.
Heterotrophic organisms, including us humans, must consume protein in order to survive. Meat is a great source, by far the best protein source per unit mass and the best for obtaining necessary protein subunits (amino acids). If you look at complete protein sources compared to caloric intake, four of the top five foods are: turkey/chicken; fish; pork chops; and lean beef.

Tofu comes in sixth and soybeans are seventh. This is why humans have sharp canine teeth – we're meat eaters. You can live happily (well, somewhat happily) as a vegetarian; you just have to work much harder at it.

So why is protein so important? How about, because it is one of the four major biomolecules and without it you die a horrible death? Sounds like a good reason to me.

Proteins reside in every cell of every living organism, from prokaryotes to your favorite uncle. There isn’t a job in a cell that proteins don’t have their hands in; proteins even perform numerous tasks at the extracellular level. Heck, that spider web hanging from your dusty Stairmaster is made of protein!

From prokaryotes to spiny echidnas to rosebushes, let’s look where proteins are involved in life. Proteins provide the structure from which cells hold their shape and onto which they build a membrane. Proteins do the talking, providing chemical signals and ways to sense chemical signals.

Proteins do the dirty work; as enzymes they put molecules together, cut them apart, and change their parts around. And most times, they make these reactions happen faster than they would otherwise and without being used up in the process.


Enzymes are specific for a very few molecules (called
substrates). Enzymes have a particular shape, and this
allows the correct substrate to bind and be acted on;
called the lock and key system. Notice that the
enzyme itself is not altered by the reaction, so it can
work again on another substrate molecule. However
there are exceptions – suicide enzymes are inactivated
by their own action, so they only work once.
Proteins allow for movement, like the contractile proteins in your muscles or the proteins that make up flagella and cilia. Proteins even act as defenders of the cell, as antibodies and myriad other immune molecules.

A typical cell may contain 10 billion protein molecules. However, not every cell has the same proteins. Many proteins are necessary for every cell, but others have specialized functions needed in only some cells. The exception is unicellular organisms. Their one cell must be able to produce every kind of protein they might ever need.

Space is at a premium, so cells can’t waste room on proteins that aren’t needed right now. Therefore, making protein must be efficient, tightly regulated, and fast. Over 2000 new protein molecules are made every second in most cells, while some proteins exist only to destroy unneeded or old proteins.

Humans can make about 2 million different proteins, but we only have about 25,000 genes that code for them. We accomplish this by having some genes produce many different proteins, just by changing the parts of the gene used. These alternative splice variant proteins may have different functions even though they come from the same gene. For example, the cSlo gene is required for hearing, and each one of the 576 different splice variants is responsible for sensing a different frequency. Biology is just so dang efficient.

Now that you know how important proteins are, let’s find out what they are. Proteins are polymers (poly = many, and mer= subunit) made up of bonded amino acid mers. Proteins come in many sizes; the TRP-Cage protein of gila monster spit is a polymer of only 20 amino acids, while the titin protein of your connective tissue is over 38,000 amino acids long.

Maybe we'll dig into the degeneracy of the genetic code when we talk about nucleic acids, but for now let’s just accept that DNA triplets code for different amino acids, and the order of the codons determines the order in which amino acids are linked to form a specific protein. The order of the different amino acids is the key. Why? I’m glad you asked.

Amino acids (or aa’s) are all small molecules made up of carbon, hydrogen, oxygen, nitrogen, and sometimes sulfur – five of last week’s “elements of life.” It’s the arrangement of these elements that makes an amino acid. Refer to the picture below for a visual aid. The central carbon is bound to four other things (often called moieities). One is simply a hydrogen. Another is an amino group (contains the nitrogen). The third is a carboxylic acid group. Get it? amino acid.


While not the most exciting images, these cartoons should
help you understand the structure of the amino acid (left)
and the building of the proteins (right). Each amino acid has
the same structure, except for whatever the R group might
be. The amino end of amino acid 2 is joined to the carboxylic
acid of amino acid 1. The next peptide bond would be between
the carboxy end of amino acid 2 and the amino end of amino
acid 3. Notice how water is created each time a peptide bond
is made.

The fourth group is what makes each aa different. Called an R group, this side chain can be small or big, neutral or charged, and gives the aa its properties. The R stands for something, but that story is just too long.

In glycine, the R group is merely another H, but in tryptophan it contains complex rings. We have talked about how tryptophan is the least used amino acid; it is bulky and introduces big bends in the peptide. We’ll show that bends, kinks and other interactions between aa’s are important for the protein function.

Most organisms can make all the amino acids they need, but mammals are the exception. We have abandoned (genetically) pathways for making some aa’s, so we must get them from our diet. These are the essential amino acids, of which there are nine if you are healthy. Tryptophan must acquired by all animals – good thing plants still have the recipe.

Ribosomes (made of proteins and nucleic acids) link the individual aa’s together in the order demanded by DNA via the mRNA. The bond that connects them is called a peptide bond, and is a “dehydration” or “condensation” reaction.

Look at the amino acid picture again; the peptide bonding process kicks out water, ie. dehydration (de = lose, and hydro – water). Water forms from seemingly nowhere, like condensation on your mirror. See how fitting the names are?

When in a protein chain (also called a peptide), the order of aa’s is called the protein’s primary (1˚) structure. The primary structure in turn dictates the secondary (2˚) structure, which is a folding of small regions of the protein based on the interactions of the side chains of closely associated amino acids.

In turn, the folding of small regions brings together aa’s from farther apart, and they fold up based on their interactions. This is the tertiary (3˚) structure of the protein. If a protein needs more than one peptide chain to be functional, the shape that those different chains form when they interact is called the quaternary (4˚) structure.

These cartoons can help you picture how an individual amino
acid can affect the structure of an entire protein. In the
secondary structure cartoon, there are two basic forms that
the nearby amino acids can form, helices and sheets, other
parts will form no patterned form at all. The tertiary and
quaternary cartoons are for hemoglobin, showing how non-
amino acids may be involved (heme), and how the
individual peptides fit together.

The hemoglobin that carries oxygen in our red blood cells is made up of four protein subunits. Why is this important – because what the protein does in life is completely dependent on its three dimensional shape. Lots of aa’s means lots of potential shapes. This is in itself one of the greatest exceptions, since one of the basic tenets of biology is “form follows function.” But with proteins, function follows form.

For the greatest number of possible combinations and shapes, it’s lucky that DNA codes for 20 aa’s. Or are there more? Proteinogenic aa’s are those that can be added into a growing peptide chain, and there are actually 22 of them. The two exceptions are selenocysteine (like cysteine with selenium substituting for sulfur) and pyrrolysine (like lysine with a ring structure added to the end).

We talked last week about the functions of selenocysteine and how it can be incorporated into a peptide even though there isn’t a normal mRNA codon dedicated to it. Pyrrolysine is similar in that it becomes coded for after the modification of what is usually a stop codon, in this case UAG (a signal to add pyrrolysine is located after the UAG codon).

Pyrrolysine is used by methanogenic (methane producing) archaea and bacteria. It's important in the active site of the enzymes that actually produce the methane. New research is showing that more organisms than previously believed use pyrrolysine. A 2011 study identified more than 16 archaea and bacteria with pyrrolysine coding mRNA modifications, but it looks like there may be more.

While the mammalian titin protein is the largest protein
known (38,136 amino acids), there is a close second in a
bacterium called Chlorobium chlorochromatii CaD3. The
gene has been found for a protein of 36,000 amino acids,
but we don’t know yet of the protein is actually made. In
archaea, the halomucin protein from the square prokaryote
Haloquadratum walsbyi is 9,200 amino acids but is exported
to protect the organism from its extreme environment.

A 2013 study indicates that the typical modification of the mRNA that occurs 100 bp downstream of the UAG stop codon isn't even there in some pyrrolysine-coding genes. One hypothesis is that in genes without the modification, the UAG sometimes acts as a stop codon and sometimes incorporates a pyrrolysine. Therefore, there are truncated (prematurely stopped) and full-length versions of the protein in the cell, and the relative number of each can be affected by local conditions and stressors.

In this paper, the authors have developed a different predictor, which doesn’t rely solely on the presence of the modification. Using it, they have identified many new candidate genes in archaea and bacteria that could be using pyrrolysines. Here’s my question – all organisms use selenocysteine, but it seems only arachaea and a few bacteria use pyrrolysine. Why did it go away in higher organisms? Can it only be used for methane production? Please, no methane production jokes.

Pyrrolysine and selenocysteine are coded for by mRNA and are added to proteins, so we definitely have 22 aa’s, but could there be more? You betcha. There are over 300 non-standard amino acids, but that isn’t such a big deal. Remember the definition of amino acid; a central carbon with a hydrogen, a carboxylic acid, an amino group, and something else attached. It isn’t a wonder there are many of them.


Bacteria kill bacteria all the time. They make their own
antibiotics, called bacteriocins, by modifying short peptides
so that they interfere with cell wall synthesis in other strains.
To do this, they modify amino acids in peptides to non-standard
amino acids, including lanthionine and 2-aminoisobutyric acid.
Those that contain lanthionine are called lantibiotics and are
hot commodities right now.
A few non-standard aa’s can be found in proteins, like carboxyglutamate which allows for better binding of calcium, and hydroxyproline, crucial in connective tissue function. These are formed by modifying the amino acids already added to the growing peptide chain.

Other non-standard aa’s are produced as intermediates in other pathways and are not used in proteins. The list of them is great and their functions are even greater, but some act as neurotransmitters, others are important in vitamin synthesis, especially in plants. Still think life uses just 20 amino acids?

Next week we can finish up proteins. Life is very selective with the form of its amino acids – except when it isn’t.


Theil Have C, Zambach S, & Christiansen H (2013). Effects of using coding potential, sequence conservation and mRNA structure conservation for predicting pyrrolysine containing genes. BMC bioinformatics, 14 PMID: 23557142

Gaston MA, Jiang R, & Krzycki JA (2011). Functional context, biosynthesis, and genetic encoding of pyrrolysine. Current opinion in microbiology, 14 (3), 342-9 PMID: 21550296


For more information or classroom activities, see:

Dietary proteins –

Functions of proteins –

Standard amino acids –

Peptide bond –

Protein structure –

Non-standard amino acids -