One easy and increasingly common way to increase the make metabolite data more useful is to associate compounds with their corresponding InChI International Chemical Identifier (InChI) (Heller et al., 2013). An InChI is a unique, standardized text representation of the structure of an organic molecule. Inclusion of InChIs in database records facilitates cross-referencing among databases. The InChI system has a number of advantages over other kinds of identifiers. Some chemical identifiers, such as PubChem IDs, Chemical Abstracts Service (CAS) numbers, and ChemSpider IDs, are database-specific accession numbers with no direct relation to the structure of the molecule they describe, this means that a molecule must have been indexed by one of these services to have an identifier. An InChI, by contrast, is a database-independent structure description, so it can be generated for a molecule regardless of whether the molecule has been indexed by a major database. An InChI can be generated for a novel natural product structure, whereas the other IDs cannot. InChI also has advantages over other linear text representations of molecule structure, such as SMILES. Unlike for SMILES, there is a single open source implementation of the InChI generation algorithm, so while a single structure may have multiple valid SMILES representations, it will only have one Standard InChI representation (Heller et al., 2013). A fixed length compressed version of an InChI, an InChIKey, can be generated from any InChI. InChIKeys are more compatible than InChIs with web search engines such as Google (google.com) (Southan, 2013), however, multiple distinct structure may have, and have been observed to have, the same InChIKey (http://www.chemconnector.com/2011/09/01/an-inchikey-collision-is-discovered-and-not-based-on-stereochemistry/), so InChIKeys should not be used as a basis for cross-referencing. When unambiguous identification of a molecule is the priority, InChI should be preferred. When ease of indexing and searchability is the priority, InChIKey should be preferred. When possible, both identifiers should be listed. By listing InChIs and InChIKeys in websites, databases, and publications (Coles et al., 2005), chemists can enhance the ability of their data to be indexed, searched, and cross-referenced. Free and easy to use software for generating InChIs and InChIKeys are the InChI software available from the InChI Trust (http://www.inchi-trust.org), and MolConverter available from ChemAxon (http://www.chemaxon.com).
Of course, even with the use of InChIs, inconsistencies can still arise in cross referencing. Galgonek and Vondrášek provide an excellent (and Open Access) analysis of the kinds of inconsistencies that can arise, and their sources.
I originally wrote this as part of a draft of the manuscript that eventually became this review article. It's a bit out of the scope of that article, so we dropped it. But I posted it here because I still think it's a good analysis.
References:
Heller et al. 2013 http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3599061/
Coles et al. 2005 http://www.ncbi.nlm.nih.gov/pubmed/15889163
Southan 2013 http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3598674/
Tuesday, October 28, 2014
Monday, October 20, 2014
What kinds of problems can (mathematical) models (of biochemical systems) solve?
I think the useful applications of computer modeling to biochemistry are in the following kind of situation: You have identified a phenotype of interest, you want to know what the molecular basis is for that phenotype, you know the form of the hypothesis (“a change in the concentration of substance X causes decreased activity of enzyme Y”), but there are too many possible instantiations of the hypothesis to test all of them experimentally (there may be 100 possibilities for “substance X” and 1000 possibilities for “enzyme Y”, giving 100,000 possible hypotheses). If you have some idea of how the system works, enough to make some kind of mathematical model of the system, and some experimental data (for example omics data comparing the condition displaying the phenotype to the condition not displaying the phenotype), you can use the computer to answer questions like: What instantiations of my hypothesis are most likely to be true, based on this dataset (or based on all available datasets)?
What is exciting?
I apologize in advance if this reads like an exercise in self indulgence. I wrote this after a conversation with a friend which basically went like this:
Her: "I'm going to go whitewater rafting next month, doesn't that sound exciting!"
Me: "Kind of. I think it would be more exciting to sit under a tree and read a good book, though."
Her: "What!? That doesn't make any sense."
It made perfect sense to me, so I wrote this essay and sent it to her. And now I'm posting it here for my throng of fans to read:
What is exciting? White water rafting? Paragliding? Sky diving? Bungee-jumping? Roller coasters? Motorcycles? These things are exciting. They are fun (at least those of these that I have experienced). But it is a short-lived excitement. It’s there, and then it’s gone.
To me, a deeper kind of excitement, one that lasts, one that I can return to again and again, comes from learning. Any idea that changes my perspective on the world or deepens my understanding is an exciting idea. It’s an idea that makes my heart beat faster.
Her: "I'm going to go whitewater rafting next month, doesn't that sound exciting!"
Me: "Kind of. I think it would be more exciting to sit under a tree and read a good book, though."
Her: "What!? That doesn't make any sense."
It made perfect sense to me, so I wrote this essay and sent it to her. And now I'm posting it here for my throng of fans to read:
What is exciting? White water rafting? Paragliding? Sky diving? Bungee-jumping? Roller coasters? Motorcycles? These things are exciting. They are fun (at least those of these that I have experienced). But it is a short-lived excitement. It’s there, and then it’s gone.
To me, a deeper kind of excitement, one that lasts, one that I can return to again and again, comes from learning. Any idea that changes my perspective on the world or deepens my understanding is an exciting idea. It’s an idea that makes my heart beat faster.
Saturday, September 13, 2014
Zamenhof's 1917 Declaration of Homaranism
In my previous post I wrote a little bit about the life of L.L. Zamenhof, and why even though he was certainly idealistic, I don't think it is fair or accurate to characterize him as "naive". In this post I present an English translation of Zamenhof's
1917 Declaration of Homaranism (Deklaracio pri Homaranismo) (pgs. 235-242 of "Mi Estas Homo", an
Esperanto anthology of Zamenhof's letters and philosophical writings). I think it is unlikely that anyone will ever establish a Homaranist organization, nevertheless, the problems that Zamenhof was trying to use Homaranism to solve are as serious today as they were in his day. So whether one agrees with the principles and methodology of Homaranism or not, I think Zamenhof's ideas on the causes and solutions to human conflict are valuable at least as food for thought for modern discussions of these same topics.
Labels:
esperanto,
heroes in their own words,
zamenhof
L.L. Zamenhof was not naive
Ludwik Lazarus Zamenhof, creator of the Esperanto language, dedicated his life to the idea of ending inter-ethnic conflict. To speakers of Esperanto (or at least to me), he's a hero. First and foremost, he's a hero to Esperanto speakers in the same way that George Lucas is to Star Wars fans, and J.S. Bach is to people who enjoy pipe organ music: he created magnificent works of art (the language, along with his translations and original writings) that continue to bring us a lot of joy. It's not an exaggeration to say that the reason Esperanto succeeded in becoming a language spoken by tens of thousands of people (or whatever the number is) all over the world was Zamenhof's stubbornness and tenacity in promoting the language and creating, from the ground up, a literature for the language. Unlike the creators of other artificial languages, Zamenhof created not only a grammar and dictionary, but powerfully demonstrated that the language was suitable even for great works of literature by translating, among many other works, examples of Shakespeare and Charles Dickens and eventually even the entire Old Testament of the Bible. Zamenhof's early works and translations were able to serve as stylistic examples to other authors and translators, and the body of Esperanto literature snowballed, and became self sustaining.
Monday, August 18, 2014
Eating really cheap: a 50 cent meal, and why I think Soylent will never be a food for the poor (although it may help them indirectly)
Two of my major interests in life are eating delicious food and spending as little money as possible. So when I read about the Soylent project, which aims to produce an ultra-cheap food powder that meets all of a person's nutrition needs I was very much interested (about the ultra-cheap part and the nutritional needs meeting part, not so much about the tasteless powder part). My philosophy about food is much different than that of Soylent creator Rob Rhinehart, who states on his blog "In my own life I resented the time, money, and effort the purchase,
preparation, consumption, and clean-up of food was consuming." Personally, I see the time I spend cooking as an adventure: an opportunity to learn and experiment. I rarely cook from recipes (although I do keep a notebook), and pretty much never eat the same thing twice. I see cooking as a kind of huge multivariable optimization problem with a complicated (and changing) objective function and no global optimum, but lots of local optima all over the place. I want to make food that is delicious, nutritious, cheap, and not monotonous, and I have fun trying to do that.
Wednesday, August 13, 2014
Chemspider python scripting
In the past, I've shown how to script Chemdraw and ChemAxon for the purposes of converting chemical information among different file formats. In this post I'll show how to achieve (vaguely) similar results using the ChemSpider web API. ChemSpider is nice because (unlike SciFinder) it is free, they don't mind if you reuse their data, and they don't mind if you access their website with scripts. Using ChemSpider has the advantage that it is a huge database and provides cross-references to other databases, it has images, mol files, and lots of other data about each compound.
Subscribe to:
Posts (Atom)
