Tampilkan postingan dengan label cheminformatics. Tampilkan semua postingan
Tampilkan postingan dengan label cheminformatics. Tampilkan semua postingan

Rabu, 06 Oktober 2010

Cheminfo Retrieval Classes 1 and 2 in 2010

My first Chemical Information Retrieval class for the Fall of 2010 took place on Sept 23, 2010. This is the second time that I've taught the class as sole instructor and it was certainly convenient to have last year's wiki to build upon. The assignments are the same so it was helpful to be able to give students access to what students did last year as examples.

The key message from my introductory lecture was that it can be really difficult to find usable chemical information and that there are no shortcuts like relying on a true trusted source - those don't exist. I showed a few examples of emerging models - Open Access, Open Notebook Science, Collaborative Competition (like pharma companies sharing some drug data openly) and other Open Science initiatives.

I also announced that we would be doing something new in the Science3.0 theme (the semantic web). One of the assignments involves collecting 5 values from the literature for each of 5 properties for a compound of the student's choice. In addition to adding these values on the wiki, we will collect them in a format that is friendly to machines: a ChemInfo Validation Google Spreadsheet. Andrew Lang has agreed to help with adapting our previous code for solubility to creating web services for this application. For example, we can have a service that reports the mean and standard deviation for a particular property and chemical. Another could produce statistics for a given data source or compare peer reviewed vs non peer reviewed sources, etc. Since it will be possible to to call these web services from within a Google Spreadsheet or Excel it should enable much more sophisticated analysis of the data related to the "validity" of chemical information as it exists today.

I didn't record the first lecture but I have the slides below:
During the second lecture on September 30, 2010 I spent most of the time showing students how to use Beilstein Crossfire, SciFinder and ChemSpider to find values for chemical properties. The recording for the second lecture is available below:

Sabtu, 02 Januari 2010

Dangerous Data: Lessons from my Cheminfo Retrieval Class

I'm not sure what my students expected before taking my Chemical Information Retrieval class this fall. My guess is that most just wanted to learn how to use databases to quickly find "facts". From what I can gather much of their education has consisted of teachers giving them "facts" to memorize and telling them which sources to trust.
Trust your textbook - don't trust Wikipedia.
Trust your encylopedia - don't trust Google.
Trust papers in peer reviewed journals - don't trust websites.
If I did my job correctly they should have learned that no sources should be trusted implicitly. Unfortunately squeezing useful information from chemistry sources is a lot of work and hopefully they learned some tools and attitudes that will prove helpful no matter how chemistry data is delivered in the future.

I have previously discussed how trust should have no part in science. It is probably one of the most insidious factors infesting the scientific process as we currently use it.

To demonstrate this, I had students find 5 different sources for properties of chemicals of their choice. Some of the results demonstrate how difficult it can be to obtain measurements with confidence.

Here are my favorite findings from this assignment as a top 3 countdown:

#3 The density of resveratrol on 3DMET

Searching for chemical property information on Google quickly reveals the plethora of databases indexed on the internet with a broken chain of provenance. These range from academic exercises of good will to company catalogs, presumably there to sell products. Although it is usually not possible to find out the source of the information, you can sometimes infer the origin by seeing identical numbers showing up in multiple places.

But sometimes the results are downright bizarre - consider the number 1.009384166 as the density of resveratrol from what looks like a Japanese government site 3DMET. First of all no units are given but lets assume this is in g/ml. The number of significant figures is curious and suggests the results of a calculation, perhaps a prediction. In this case the source is from the MOE software. This is clearly a different algorithm from the one used by ACDLabs, which comes in at 1.356 g/ml, much more realistic when put up against all 5 sources:
  • 1.359 g/cm3 ChemSpider predicted
  • 1.36 g/cm3 (20 C) Chemical Book MSDS
  • 1.009384166 3DMed
  • 1.41 g/cm3 (-30.15C) DOI (found with the aid of Beilstein)
  • 1.359 g/cm3 LookChem
#2 The melting point for DMT depends on the language

I have to admit being really surprised by this. Even though I knew that Wikipedia pages in different languages were not exact translations I would have assumed that the chemical infoboxes would not be recreated. Interestingly, the German edition has a reference but I was not able to access it since it is a commercial database. The English edition has no specific references. Here is a list of sources:
#1 Solubility of EGCG in water

This is by far my favorite because it most clearly demonstrates the dangers of the concept of a "trusted source". From the compilation prepared by the student, this paper (Kwang08) reported the solubility of EGCG at 521.7 g/l:

This is from a paper that spent 5 months undergoing peer review with a well respected publisher. Also it appeared recently so one would expect the benefit of the best instruments and comparison with historical values. But even beyond all of this, the numbers are in the opposite order to the point explained in the paragraph. In our system of peer review we don't expect reviewers to verify every data point - but we do expect the text to be evaluated as logically consistent.

Now if we follow the reference provided for this paragraph we find the following paper (Liang06), with this:

We can now see what happened: the 21.7 was accidentally duplicated from the caffeine measurement and appended to the 5 g/l for EGCG. This is a lot more reasonable, even though I am not clear about where that number comes from in this second paper.

We can get some idea of the potential source of this information from the Specification Sheet for EGCG on Sigma-Aldrich:

Notice that this does not state that the maximum solubility of EGCG in water is 5 mg/ml - just that a solution of that concentration can be made. This value is repeated elsewhere, such as this NCI document, which references Sigma-Aldrich:
From here the situation gets muddled. Another search reveals this peer reviewed paper (Moon06), which appeared in 2006:
Expressed in mM this translates to about 2.3 g/l. Clearly this value is inconsistent with the Sigma-Aldrich report of being able to make a clear solution at 5 g/l.

Luckily, in this case we have some details of the experiments:

The measurements were done in triplicate and averaged. Unfortunately this does not reveal any sources of systematic error. One clue as to why these values are contradictory might be the method of dissolution. One hour sonication at room temperature might just not be enough to make a saturated solution for this compound. (Although one might expect the error to lie on the high side because the sample were diluted before being filtered) What would answer this definitively are the experimental details of how the Sigma-Aldrich source prepared the 5 g/l solution. If it went in within a few minutes without much agitation, that would be inconsistent with this hypothesis of insufficient mixing. In that case we would want to look at the HPLC traces in this paper for another type of systematic error.

Unfortunately, the chain of information provenance ends here. Just based on the data provided so far, there is significant uncertainty in the aqueous solubility of EGCG, similar to our uncertainty about the melting point of strychnine.

As long as scientists don't provide - and are not required to provide by publishers - the full experimental details recorded in their lab notebooks, this type of uncertainty will continue to plague science and make the communication of knowledge much more difficult than it need be.

Unfortunately the concept of "trusted sources" is being used as a building block of some major chemical information projects currently underway - WolframAlpha and the chemical infobox data of Wikipedia are prime examples. Ironically, MSDS sheets are listed as a reliable "trusted source" for the infoboxes, when they have been shown to be very unreliable (see my previous post about this with statistics). These are probably one of the most dangerous sources of information because they appear to be trustworthy - coming from chemical companies and the government - and often found on university websites. Combine that with the absence of references or experimental details and the potential for replication of errors is very high and very difficult to correct.

WolframAlpha does have a mechanism to provide information about sources but it requires submitting a reason and personal information.
To see how this works in practice I made a request for the source of an entry with erroneous data - glatiramer acetate:
I submitted this 10 days ago and still don't know the source.

Rapid access to specific sources is important for maximizing the usefulness of databases. Without that it becomes very difficult to assess the meaning of reported measurements and compare with results from other databases.

It is not possible to remove all errors from scientific publication. But that's only a problem when it is difficult to determine that there are errors in the first place because insufficient information is provided.

Scientists can handle ambiguity. If you look at the discussion over the blogosphere concerning the JACS NaH oxidation paper, much of it was constructive. The publication of that paper was not a failure of science. Quite the opposite - we learned some valuable lessons about handling this reagent. As far as I can tell the paper was a truthful reporting of their results.

Where this was a failure lies in the way conventional scientific channels handled the matter. There was no mechanism to comment directly on the website where paper was posted. That would have been the logical place for the community to ask questions and have the authors respond. Instead the paper was withdrawn without explanation.

Kamis, 05 November 2009

Sixth Cheminfo Retrieval class: What is the m.p. of strychnine?

It would seem to be a simple task to find the melting point of a well known alkaloid like strychnine. Our quest to answer that question - and other simple properties - in class using both freely available and commercial databases reveals how treacherous it can be. In the end we don't find an unambiguous answer but we uncover enough information for many applications.

The take home message is that chemists need to be constantly paranoid that their information - whether from their lab or the most prestigious journals - can easily be wrong. Strategies such as finding multiple sources and investigating the experimental details provided in the primary sources are demonstrated to diminish uncertainty. But this is often not easy or quick.

Here is a summary of the lecture:

This is the lecture from the sixth Chemical Information Retrieval class at Drexel University on October 29, 2009. It starts with a review of some of the new questions answered by students from the chemistry publishing FAQ, which covers patent information and accessing electronic journals at Drexel. Tony Williams submitted a puzzle to resolve conflicting structures in ChemSpider, which is too difficult to be a regular assignment. It requires re-analyzing spectroscopic data in papers where stereochemical assignments are determined. An example is paromomycin which has three entries. The regular assignment for the week is then introduced and it involves obtaining 5 different sources each for 5 different properties for a molecule of the student's choosing. To demonstrate how to do the assignment strychnine is chosen as an example. Melting point information is obtained from ChemSpider (ultimately an MSDS sheet), Wikipedia, Wolfram Alpha and in a JACS article via SciFinder. By investigating primary sources several errors are found in SciFinder, where the recorded melting points correspond to salts of the alkaloid. Difficulties in finding primary sources for the melting point from Wikipedia are highlighted. For LD50 information Wikipedia did not even provide proper units (mg instead of mg/kg and no animal or route specified). The importance of ChemSpider predicted values for density and boiling point is demonstrated as a corroborating tool. In the end the reported melting point range of strychnine from the JACS paper did not even overlap with the reference to which it was compared. The exercise is meant to highlight the importance of caution in obtaining values from all available sources. Even the seemingly simple question of determining the melting point of well known alkaloid cannot be answered definitively.

Rabu, 04 November 2009

Glatiramer Acetate Cheminformatics Problem and Fifth ChemInfo Retrieval Class

It started out innocently enough. One of my students picked the multiple sclerosis drug glatiramer acetate for his project in my Chemical Information Retrieval class. This ultimately resulted in the removal of this substance from ChemSpider.

The problem is that this drug is a polymer but it is represented in many places as a simple mixture of acetic acid and 4 amino acids (L-Ala, L-Glu, L-Lys, and L-Tyr). See for example Wikipedia, PubChem and DrugBank.


The SMILES representation is entered as 5 molecules joined by periods:
CC(O)=O.C[C@H](N)C(O)=O.NCCCC[C@H](N)C(O)=O.N[C@@H](CCC(O)=O)C(O)=O.N[C@@H](CC1=CC=C(O)C=C1)C(O)=O
This is probably the source of all subsequent miscalculations - such as a molecular weight of 623.7 (it actually has an average MW one order of magnitude larger), molecular formula C25H45N5O13, Topological Polar Surface Area of 374, Rotatable Bond Count 13, a 3D structure that is nowhere near reality, etc.

Glatiramer acetate is reported to bind to MHC molecules. If these molecular descriptors are used in any type of QSAR analysis this will just add noise to the models.

ChemSpider does not keep track of polymers, except perhaps for some well defined oligopeptides that can be represented by a single SMILES. Consequently it was removed from the database.

It is difficult to apply common cheminformatics tools to this substance. It might be tempting to try to place it in polypeptide/protein databases such as BioPD. But it does not have a well defined length or composition. In fact it is a random co-polymer so it can not even be represented by a repeating structure, such as one might do for polystyrene.

In order to generate meaningful molecular descriptors for QSAR applications I suppose one strategy would be to generate a collection of SMILES representing the average composition of the drug in terms of ratios of amino acids and molecular weights. Each structure would generate molecular descriptors and 3D structures that are far more realistic than those currently listed. Perhaps it would turn out that only some of these polymer structures interact with MHC molecules. (If this has already been done please forgive the oversight - I didn't research this thoroughly. By the end of the term we should know more from the student's report)

The chronological summary of the lecture is as follows:

The fifth Chemical Information Retrieval class on October 22, 2009 started out with covering the new 3D structure viewer introduced recently at PLoS ONE to provide ideas for students doing a multimedia project this term. The current student answers to the chemistry publishing FAQ are then discussed. The reason for removing glatiramer acetate from ChemSpider is explained and a few databases (Wikipedia, PubChem, DrugBank) are visited that still contain the incorrect SMILES, 3D structure and related properties. An overview of an Open Access site (OAD) suggested by Bill Hooker is provided to suggest additional questions for the FAQ. Examples of questions discussed include primary and secondary sources, peer review, article level metrics (a PLoS ONE article on malaria is used as an example), citation searching, Impact Factors and whether one should use one's real name in the blogosphere. Databases Scirus, Web of Science and PubMed are also reviewed.

Selasa, 20 Oktober 2009

Fourth Cheminfo Retrieval class: ChemSpider and Beilstein Databases

Peggy Dominy, our chemistry librarian at Drexel, was kind enough to teach my third class while I was at NERM. She demonstrated RefWorks - including how to copy and paste the proper formats to Wikispaces - and how to use our ILL (Inter-Library Loan) process.

I'm including a recording of the fourth class on Chemical Information Retrieval on Oct 15, 2009 at Drexel University. It starts with some tips on removing formatting from Wikispaces pages, the Drexel Cisco VPN client for accessing paid subscriptions off campus and how to link to a DOI. The first two assignments for the class are then described. The first involves summarizing each paragraph of an article and an option to use AcaWiki is demonstrated. The second involves filling in an FAQ for publishing in chemistry. FriendFeed is then presented as a resource to help answer questions followed by an extensive overview of available information on ChemSpider, covering SMILES, InChIs, InChIKeys, experimental and predicted properties, linked databases and contributed spectra. Finally a demonstration of Beilstein Crossfire/DiscoveryGate is presented with an emphasis on doing substructure searching.

Jumat, 25 September 2009

Cheminfo Retrieval First Class FA09

I gave my first lecture yesterday (Sept 24, 2009) for my Chemical Information Retrieval course at Drexel. One of my main objectives for the course is to provide the most current information about how to best find and review chemical information.

To this end, I set up a wiki (http://getcheminfo.wikispaces.com) which should become considerably enriched over the course of the term. I invited students to help contribute useful links to the resource page - and even before I finished giving the first lecture they added several really good ones. I also invite any chemists or librarians to add links to resources we may have missed. Just request to join the wiki to contribute.

The wiki will also be used for students to write a report on a chemical topic making use of cheminfo resources. Right after the lecture I made sure the students joined the wiki and created two pages: one for their report and one for a "research log". The idea is that students will report significant steps in conceptualizing their projects and how they are searching databases. I can then comment directly on their log pages for quick guidance. I suppose anyone with helpful suggestions that I missed could also comment - again just request an account on the wiki.

This class has traditionally required a written report. This term I'm adding a twist: the minimum number of words can be reduced somewhat if students elect to incorporate a multimedia or other creative component. To provide examples of what that might look like I visited Drexel Island on Second Life and demonstrated 3D molecules, interactive NMR spectra and a chemistry museum (from Sandy Adam). There is a lot of chemistry possible on Second Life (see Lang & Bradley) At the end of the tour on the island we visited a wildlife area recently built by Robert Brulle for a project related to environmental science (more on this in a later post). I got a hug from a panda and got sprayed when I tried to pet a skunk - just to give a taste of what kind of fun things can be constructed in a virtual world. Other projects could involve screencasts, Jmol, games, Facebook, etc. As long as it requires students to access chemical information, I am pretty open to ideas. Students will work through their ideas on their log page and the final product will also be available on the wiki. These projects could provide interesting examples for others interested in the topic of chemical information.

At the end of my lecture I provided a brief overview of the NaH oxidation controversy. There really could not be a better example of the importance of staying on top of new communication channels to follow and participate in chemical research. This year the most important of these new tools are probably blogs, wikis and FriendFeed. Next year it might be something else - Google Wave?


Rabu, 18 Maret 2009

Chemistry publication - making the revolution

Steven Bachrach has just published a very interesting commentary "Chemistry publication - making the revolution" on March 17, 2009 in the new Journal of Cheminformatics (I am on the editorial board).

The article does a great job of highlighting the current state of affairs in chemistry which is creating a chasm between researchers and the data they could be using. There are suggestions of steps that can be taken by researchers, reviewers and editors. Although Open Notebook Science is mentioned, there are several less intensive steps that can be taken to still benefit the community. Steven mentions the submission of JCAMP-DX files instead of PDFs when submitting spectra as supplementary materials for publications. I agree that this is a very low barrier step that chemists can take to gain immediate benefit - whether they share their spectra or not. At that point it also becomes trivial to contribute spectra to databases like ChemSpider that supports the open JCAMP-DX format - hopefully as Open Data (so that it can be used in applications such as our SpectralGame)

Jumat, 17 Oktober 2008

Journal of Cheminformatics

Set to come out within the next few months, the Journal of Cheminformatics has it all for the Open Access hungry mob:

Editor-in-Chief
Dr David J Wild, Assistant Professor of Informatics at Indiana University, USA

Chemistry Central announces the imminent launch of the Journal of Cheminformatics.

The Journal of Cheminformatics is devoted to the dissemination of new and original knowledge in all branches of cheminformatics and molecular modelling including:

  • chemical information systems, software and databases, and molecular modelling
  • chemical structure representations and their use in structure, substructure, and similarity searching of chemical substance and chemical reaction databases
  • computer and molecular graphics, computer-aided molecular design, expert systems, QSAR, and data mining techniques

Chemistry Central publishes peer-reviewed open access research in chemistry.

All research articles published by Chemistry Central are made freely and permanently accessible online immediately upon publication to ensure effective communication of research findings.

High standards are maintained through full and stringent peer review.

Authors who publish original research in Chemistry Central retain copyright over their work.

Chemistry Central is an independent publishing service operated by BioMed Central - the leading life science open access publisher.

Jumat, 01 Agustus 2008

The BCCE, research discussions and good friends

I've just returned from the Biennial Conference on Chemical Education 2008.

I wasn't able to record my first talk on Communicating Results from Undergraduate Research because my computer crashed. However, I repeated many of the same concepts in my second talk on Open Notebook Science and Cheminformatics.

It was very fortunate that the BCCE was held in Bloomington at Indiana University this year because it was a great opportunity for me to meet up with Rajarshi Guha, David Wild and Amar Flood.

Amar and I discussed Open Notebook Science with his lab people and it may make sense to do this for some of their projects. We set up a wiki to explore that possibility.

Rajarshi and I discussed at length our collaboration on the prediction of Ugi precipitates and docking against falcipain-2. It is certainly easier to pour over the relevant papers and online documents when face to face.

These are the outcomes:

1) We're going to separate the problem of predicting the solubility of Ugi products from the problem of generating the best enzyme inhibitors. Our initial plan was to try to make the top ranked Ugi products for a given enzyme and hope that we generate enough precipitates from those results for Rajarshi to model. We're just not getting enough positive results that way to generate a reliable model in a reasonable amount of time. To stack the deck in our favor we're going to start with a well behaved Ugi reaction (EXP099) and modify the reagents one at a time.

2) We're going to separate the performance of the Ugi reaction from the solubility of Ugi products in various solvents. Once we have Ugi products in hand in pure form from a reaction in methanol we will simply measure their solubility in other solvents, starting with ethanol, acetonitrile, THF and toluene. Low solubility will not guarantee that the Ugi reaction will proceed smoothly to produce a precipitate when carried out in a given solvent but it will certainly be a great starting point.

3) We're going to actually measure the solubility instead of just noting soluble or insoluble, as we have been doing in our Ugi master table. This will make it much easier for Rajarshi to come up with a robust model. Kevin Owens is already set up with a SpeedVac in his lab and that will help immensely. Rajarshi made the point that models predicting solubility in non-aqueous systems are needed and could be quite helpful to the chemistry community. He will be using the crystallographic data of our precipitates in his calculations. More on this later...

4) We're going to require the reactions to be easily amenable to automation, even if we can carry them out manually sometimes. For example, some of the reactions had starting materials that were not very soluble in methanol by themselves, even though they went into solution when combined with the other starting materials and then generated a product. This is interesting behavior but extremely inconvenient for automation because we can't make up stock solutions of reagents at 2M concentration in methanol.

5) We're going to require the reactions to be fast. Some reactions required several days to complete. These will now be considered to be negative for precipitation at the 16 hour mark.

Minggu, 27 April 2008

UsefulChem on ChemSpider

As Antony Williams mentions in his blog, UsefulChem molecules now have their own url on ChemSpider:
http://usefulchem.chemspider.com

This is important on so many levels. I've been talking about using ChemSpider as an integral part of our research for some time now and we've been steadily migrating content as new functionalities became available. We still have a lot to do to complete the process but the path is clear.

The ability to upload spectra of all types (IR, H NMR, C NMR, MS, etc) in JCAMP-DX format allows for anyone with a browser to drill down to any desired expansion using JSpecView. It just makes sense to upload the NMRs of our starting materials and our fully characterized products so that we can always be just a click away from the online lab notebook when discussing our chemistry.

The value of being able do substructure searching for our compounds is immense. Most organic chemistry students are not interested in learning how to run and maintain software on a lab server to keep track of their molecules. They also don't want to learn SMARTS. ChemSpider has a very familiar graphic interface for drawing molecule fragments to do substructure searching. For a free and hosted solution, the energy of activation for using ChemSpider has become very low indeed, especially for Open Notebook Science applications.

As we approach 200 experiments it is becoming clear that the ability to retrieve information is just as important as doing the experiments. Systematic tagging of experiment pages with InChIKeys and InChIs is a simple way to automatically allow for direct links from ChemSpider to UsefulChem pages via Google.

We've come a long way from using a blog to track our molecules :)

Kamis, 10 April 2008

Back from Spring08 ACS

The American Chemical Society meeting in New Orleans proved to be as hectic as usual. I didn't get to spend as much time with everyone I was hoping to touch base with but I did get some useful things done:

1) I presented on the use of Cheminformatics in Open Notebook Science and Teaching Chemistry with Second Life. Since the ACS oral sessions are so short (15-20 mins before questions) this is a good opportunity to get our work recorded in a concise format. Speakers like to complain that they need a full hour to do justice to their elaborate projects but the shorter format seems to match the attention span of most attendees.

2) The Sci-Mix session on Monday was very hectic. I presented two posters related to my talks and helped out with the opening of the poster session on ACS island. We had 20 posters in Second Life and it was great to finally meet a few of the presenters and the organizer Kate Sellar (Finola Graves) in person. Kate was running Second Life on a large screen right at the poster entrance and giving out t-shirts. Andy Lang (Hiro Sheridan) was also helping out remotely from within Second Life. By all accounts it was well received. The posters will remain up on ACS island for the next few months and it will be possible to contact presenters via bells next to their posters.

3) I also finally met several people that I have known only virtually for some time: Rajarshi Guha, David Wild, Rich Apodaca, Noel O'Boyle and Warren DeLano. Rajarshi, Rich, Noel, Warren, Geoff Hutchison and I went out for the Blue Obelisk dinner on Tuesday night.

4) I had a good brainstorming session with Rajarshi about representing experimental workflows in XML. He suggested using Taverna to create "null" workflows - since our lab experiments are physical rather than executing code. Noel suggested using Knime. The idea is to represent experimental steps leading to a result in a standard machine readable (and thus minable) format. We're going to start very simply by limiting the workflow space to Ugi reactions.