<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://chem-bla-ics.linkedchemistry.info/feed/by_tag/cml.xml" rel="self" type="application/atom+xml" /><link href="https://chem-bla-ics.linkedchemistry.info/" rel="alternate" type="text/html" /><updated>2026-07-18T13:36:15+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/feed/by_tag/cml.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">The Molecular Chemometrics Principles #2: be clear in what you mean</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html" rel="alternate" type="text/html" title="The Molecular Chemometrics Principles #2: be clear in what you mean" /><published>2010-08-12T00:00:00+00:00</published><updated>2010-08-12T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html"><![CDATA[<p>I noted <a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">earlier this week</a>
that <em>[d]uring the week [in <a href="/2010/08/06/oxford-2.html">Oxford <i class="fa-solid fa-recycle fa-xs"></i></a>], someone (name and address is know at the
editorial office) commented on the fact that my blog posts are somewhat difficult to follow; that is, it’s
often not clear why I am posting what I am posting</em>. This triggered the start of a series of principles in
the field I coined <a href="https://doi.org/10.1080/10408340600969601">Molecular Chemometrics</a>, and the promise
that I will try to indicate in each blog post to which of these principles it relates. Just to put things in a bit more
perspective; to make a bit more clear why I am blogging about that bit; just to be clear in what I mean.</p>

<p>Now, the first principle was about the need for access to data (<a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">McPrinciple #1</a>).
This principle goes without saying, one would think, but is not widely accepted yet. This is why Open Data promotion is still needed. For example, data in papers
still is not freely redistributable, as <a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">Peter points out once again</a>.</p>

<p>Anyway, this post is not about McPrinciple #1, but about the second principle.</p>

<p><strong>Molecular Chemometrics Principles #2</strong>: In order to reproduce cheminformatics studies you need to be able to understand the input data.</p>

<p>Readers of my blog will surely recognize this theme. Clearly this theme explains my past fetish for the
<a href="http://chem-bla-ics.blogspot.com/search?q=CML">Chemical Markup Language</a>, and my more recent work on the
<a href="http://chem-bla-ics.blogspot.com/search?q=RDF">Resource Description Framework</a>.</p>

<p>And it is so easy to jump to conclusions. Easy to make mistakes. And this is not just at the received side; the sending
person may have accidentally made a mistake, or left something accidentally unclear, causing incorrect assumptions, and
therefore errors in the cheminformatics computation. Now, if the data was semantically (clearly) annotated, and the
meaning was clear, it was also trivial to see when a mistake had sneaked in. Think of it as a check bit.</p>

<p>“Well, isn’t this a bit exaggerated,” you might say. Perhaps, perhaps not. An simple, recent example. We all know
<a href="http://www.opensmiles.org/">SMILES</a>, right? And we all know that lower case element symbols indicate aromaticity, right?
That is, c1ccccc1 is aromatic, right? So, what’s the problem then?</p>

<p>Now, consider the SMILES string c1ccc1. Lower case carbon element symbols, so aromatic, right? Oh, wait…</p>

<p>Therefore, be clear in what you mean. It saves us from a lot of trouble.</p>

<p>Further reading:</p>

<ul>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">The Molecular Chemometrics Principles #1: access to data</a></li>
  <li>Molecular Chemometrics, 2006 (doi:<a href="https://doi.org/10.1080/10408340600969601">10.1080/10408340600969601</a>)</li>
</ul>]]></content><author><name>Egon Willighagen</name></author><category term="mcprinciples" /><category term="chemometrics" /><category term="rdf" /><category term="cml" /><category term="semweb" /><category term="doi:10.1080/10408340600969601" /><summary type="html"><![CDATA[I noted earlier this week that [d]uring the week [in Oxford ], someone (name and address is know at the editorial office) commented on the fact that my blog posts are somewhat difficult to follow; that is, it’s often not clear why I am posting what I am posting. This triggered the start of a series of principles in the field I coined Molecular Chemometrics, and the promise that I will try to indicate in each blog post to which of these principles it relates. Just to put things in a bit more perspective; to make a bit more clear why I am blogging about that bit; just to be clear in what I mean.]]></summary></entry><entry><title type="html">Extracting RDF from Chem4Word documents</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/01/21/extracting-rdf-from-chem4word-documents.html" rel="alternate" type="text/html" title="Extracting RDF from Chem4Word documents" /><published>2010-01-21T00:00:00+00:00</published><updated>2010-01-21T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/01/21/extracting-rdf-from-chem4word-documents</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/01/21/extracting-rdf-from-chem4word-documents.html"><![CDATA[<p><a href="http://jat45.wordpress.com/">Joe</a> has released the first <a href="http://research.microsoft.com/en-us/projects/chem4word/">Chem4Word</a>
<a href="http://jat45.files.wordpress.com/2010/01/example.docx">demo file</a>, and has written about how to
<a href="http://jat45.wordpress.com/2010/01/20/extracting-cml-from-a-chem4word-authored-document-java/">extract the CML with Java</a>
and <a href="http://jat45.wordpress.com/2010/01/21/extracting-cml-from-a-chem4word-authored-document-c/">with C#</a>.</p>

<p>I haven’t actually gotten around to fiddling with Java, but ran <a href="http://strigi.sf.net/">Strigi</a> against it to extract RDF,
while having the <a href="http://neksa.blogspot.com/2007/05/introduction.html">Strigi-Chemistry</a> plugins installed. This is part of the
<a href="http://en.wikipedia.org/wiki/Resource_Description_Framework">RDF</a> that came out:</p>

<div class="language-turtle highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nl">&lt;example-doc.docx&gt;</span><span class="w">
  </span><span class="nl">&lt;http://freedesktop.org/standards/xesam/1.0/core#title&gt;</span><span class="w">
    </span><span class="s">"acetic acid"</span><span class="p">,</span><span class="w">
    </span><span class="s">"(8R,9S,10R,13S,14S,17S)- 17-hydroxy-10,13-dimethyl- 1,2,6,7,8,9,11,12,14,15,16,17-dodecahydrocyclopenta[a] phenanthren-3-one"</span><span class="p">,</span><span class="w">
    </span><span class="s">"testosterone"</span><span class="p">;</span><span class="w">
  </span><span class="nl">&lt;http://freedesktop.org/standards/xesam/1.0/core#version&gt;</span><span class="w">
    </span><span class="s">"2"</span><span class="p">,</span><span class="w">
    </span><span class="s">"2"</span><span class="p">;</span><span class="w">
  </span><span class="nl">&lt;http://rdf.openmolecules.net/0.9#atomCount&gt;</span><span class="w">
    </span><span class="s">"8"</span><span class="p">,</span><span class="w">
    </span><span class="s">"49"</span><span class="p">;</span><span class="w">
  </span><span class="nl">&lt;http://rdf.openmolecules.net/0.9#bondCount&gt;</span><span class="w">
    </span><span class="s">"7"</span><span class="p">,</span><span class="w">
    </span><span class="s">"52"</span><span class="p">;</span><span class="w">
  </span><span class="nl">&lt;http://rdf.openmolecules.net/0.9#molecularFormula&gt;</span><span class="w">
    </span><span class="s">"C2H4O2"</span><span class="p">,</span><span class="w">
    </span><span class="s">"C19H28O2"</span><span class="p">;</span><span class="w">
</span></code></pre></div></div>

<p>I believe there is quite some room for improvement, but it’s a start :) Thanx to Joe for posting the public domain test file, so
that other projects can start play with the exiting new technology. I should note, however, that I am not running a Microsoft OS
nor MS-Word, and the saved documents source are the only way I have access to the
<a href="http://en.wikipedia.org/wiki/Chemical_Markup_Language">CML</a> right now.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cml" /><category term="java" /><category term="rdf" /><category term="chem4word" /><category term="strigi" /><summary type="html"><![CDATA[Joe has released the first Chem4Word demo file, and has written about how to extract the CML with Java and with C#.]]></summary></entry><entry><title type="html">CrossRef writes up RSS usage recommendations</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/10/20/crossref-writes-up-rss-usage.html" rel="alternate" type="text/html" title="CrossRef writes up RSS usage recommendations" /><published>2009-10-20T00:00:00+00:00</published><updated>2009-10-20T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/10/20/crossref-writes-up-rss-usage</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/10/20/crossref-writes-up-rss-usage.html"><![CDATA[<p><a href="http://www.crossref.org/CrossTech/2009/10/recommendations_on_rss_feeds_f.html">CrossTech announced</a> that a <a href="http://www.crossref.org/">CrossRef</a>
working group has written a <a href="http://oxford.crossref.org/best_practice/rss/">best practices</a> for the use of RSS feeds by publishers. Nice introduction
for anyone who is creating RSS feeds. Only comment I could make, is the lack of other modules. For example, a Chemistry module has been proposed by
us 5 years ago already (DOI:<a href="http://dx.doi.org/10.1021/ci034244p">10.1021/ci034244p</a>) and about which I blogged on
<a href="http://chem-bla-ics.blogspot.com/search?q=CMLRSS">several occasions</a>.</p>

<p>Below is the <a href="http://cb.openmolecules.net/atom.php?category=&amp;type=latest_inchis">CMLRSS feed</a> of <a href="http://cb.openmolecules.net/">Chemical blogspace</a>.</p>

<p><img src="/assets/images/cmlrss_Cb2.png" alt="" /></p>

<p>Of course, publishers can take advantage of such modules, using the <a href="http://www.w3.org/TR/xml-names/">XML Namespaces</a> technology. The <em>best practices</em>
uses that for a <a href="http://dublincore.org/">Dublin Core</a> and a <a href="http://purl.org/rss/1.0/modules/prism/">PRISM</a> extension. The here discussed CML
extension is another one, but the point is, that you can basically plug in any module.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cml" /><category term="rss" /><category term="xml" /><category term="doi:10.1021/CI034244P" /><summary type="html"><![CDATA[CrossTech announced that a CrossRef working group has written a best practices for the use of RSS feeds by publishers. Nice introduction for anyone who is creating RSS feeds. Only comment I could make, is the lack of other modules. For example, a Chemistry module has been proposed by us 5 years ago already (DOI:10.1021/ci034244p) and about which I blogged on several occasions.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cmlrss_Cb2.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cmlrss_Cb2.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Journal of Cheminformatics: I hope the Instructions to the Authors improve</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/03/22/journal-of-cheminformatics-i-hope.html" rel="alternate" type="text/html" title="Journal of Cheminformatics: I hope the Instructions to the Authors improve" /><published>2009-03-22T00:00:00+00:00</published><updated>2009-03-22T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/03/22/journal-of-cheminformatics-i-hope</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/03/22/journal-of-cheminformatics-i-hope.html"><![CDATA[<p>Besides <a href="https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html">Nature Chemistry <i class="fa-solid fa-recycle fa-xs"></i></a>, another journal was launched last week (see
<a href="http://www.steinbeck-molecular.de/steinblog/index.php/2009/03/17/open-access-journal-of-cheminformatics-now-live/">here</a> and
<a href="http://blogs.openaccesscentral.com/blogs/ccblog/entry/journal_of_cheminformatics_publishes_launch">here</a>): the
<a href="http://www.jcheminf.com/">Journal of Cheminformatics</a>. First of all, congratulations to <a href="http://www.steinbeck-molecular.de/steinblog/">Chris</a>
and David for their efforts! While the journal only published one research paper yet, it already found
<a href="http://cb.openmolecules.net/journal_search.php?journal_id=Journal%20of%20Cheminformatics">its place</a> on
<a href="http://cb.openmolecules.net/">Chemical blogspace</a>. I have two things I want to blog about: <em>data rich publishing</em>, and
<em>starting the scientific communication</em>.</p>

<h2 id="data-rich-publishing">Data Rich Publishing</h2>

<p>Peter had a <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=1326">detailed blog</a> about why he joined the editorial board:</p>

<blockquote>
  <p>I take this position with some trepidation as I have grave reservations about the current practice of cheminformatics.
It suffers from closed data, closed source and closed standards, and thereby generally poor experimental design, poor
metrics and almost always irreproducible results and conclusions which are based on subjective opinions.</p>
</blockquote>

<p>I strongly agree with this observation, and have discussed my view on this in
<a href="https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html">my thesis <i class="fa-solid fa-recycle fa-xs"></i></a> (send me an email if you
want a copy).</p>

<p>So, what has the journal to say about this (see <a href="http://www.jcheminf.com/info/instructions/">Instructions to the Author</a>,
emphasis mine):</p>

<blockquote>
  <p>Journal of Cheminformatics recommends, <strong>but does not require</strong>, that the source code of the software should be made
available under a suitable open-source license that will entitle other researchers to further develop and extend
the software if they wish to do so.</p>
</blockquote>

<p>Regarding data, they even less revolutionary; recommended figures formats (EPS, PDF, PNG) focus on nice graphics instead
of reuse of data. I also note that I cannot upload data in the <a href="http://en.wikipedia.org/wiki/OpenDocument">Open Document Format</a>,
or, hey, let’s really push things, in <a href="http://en.wikipedia.org/wiki/Resource_Description_Framework">RDF</a>. Well, not according to
the Instructions. And surely, I can put the [O|R]DF in the supplementary information, anyway. It would also be nice if I could
use Jmol as an applet to enrich the graphics, and improve data reusability of the paper, like the
<a href="https://chem-bla-ics.linkedchemistry.info/2009/01/19/rsc-now-allows-jmol-in-main-text-of.html">RSC recently started to allow <i class="fa-solid fa-recycle fa-xs"></i></a>.</p>

<p>Regarding the supplementary information, there is a section on <em>additional files</em>, which, unconveniently are capped at
20MB size. No mention of chemical formats at all, neither any recommendation on semantic formats like
<a href="http://en.wikipedia.org/wiki/Chemical_Markup_Language">CML</a> (I wonder when this was discussed with the Editorial Board,
and where Peter was at the time). How am I going to put online my 500 molecular structure CML file now? (Though it’s good
to know it is virus scanned ;)</p>

<p>So, why do I vent my concerns about these limitations? I had not blogged about the launch of the journal earlier, because
I have not made up my mind about it. On one side, I am happy to see a journal that promotes (scientific) use of papers,
and a journal that allows me to keep copyright on the material. However, on the other side, what the current Instructions
suggest, the data I could use from the papers is available only in an old-fashion way. That’s a lost opportunity and could
have killed competition for sure. Instead, the unique selling point is now restricted to using an
<a href="http://www.biomedcentral.com/info/about/openaccess/">open access license</a>. Nature Chemistry, on the other hand, chose
data rich publishing as a selling point (though in competition with things done at the RSC).</p>

<p>The other thing I want to mention about the journal is the following. <a href="http://blog.rguha.net/">Rajarshi</a> blogged about
<a href="http://hackberry.chem.trinity.edu/blog/">Bachrach</a>’s paper on <em>Chemistry publication - making the revolution</em>
(DOI:<a href="https://doi.org/10.1186/1758-2946-1-2">10.1186/1758-2946-1-2</a>). Firstly, by adding a link like that for the
DOI I just gave, Chemical blogspace can pick it up; we need this later. Secondly, the paper actually suggests that
<em>“[b]y publishing lots of data, available for ready re-use by all scientists, we can radically change the way science
is communicated and ultimately performed”</em>; this is in strong contrast to what I have seen in the Instructions so far.</p>

<h2 id="starting-the-scientific-communication">Starting the Scientific Communication</h2>
<p><a href="http://depth-first.com/">Rich</a> <a href="http://blog.rguha.net/?p=216#comment-342">replied</a> to Rajarshi about the requirement
to log in before someone could make a comment, which he did not like. He suggested alternative ways to prevent SPAM
and sorts. The choice for this commenting approach may also originate from having an Open discussion, where everyone
takes responsibility for what he says. The use of OpenID, as Rich suggests would only partially address that; on the
other hand, setting up a fake email address is quite common in the blogosphere too.</p>

<p>If Rajarshi would have used the DOI to link to the Steven’s paper, as said, Chemical blogspace would have recognized
it. Instead, he chose to link directly to the PDF. This is a typical case of hamburgers in action. However, others
did when they discussed the first research paper in the journal (DOI:<a href="https://doi.org/10.1186/1758-2946-1-3">10.1186/1758-2946-1-3</a>).
These blogs were picked up by Cb and are listed on <a href="http://cb.openmolecules.net/paper.php?paper_id=1666">this page</a>.</p>

<p>Now, I only need to remind you of <em>Userscripts for the Life Sciences</em> (DOI:<a href="https://doi.org/10.1186/1471-2105-8-487">10.1186/1471-2105-8-487</a>)
that we have the methods to link these comments back to the journal website. The <em>Quotes from Chemical Blogspace and Postgenomic</em>
script in particular, does the hard work (needs GreaseMonkey, the script can be downloaded here; see also
<a href="http://baoilleach.blogspot.com/2007/04/add-quotes-from-postgenomic-and.html">Noel’s original post</a>). This way,
we can read the comments when we visit the <a href="http://www.jcheminf.com/content/1/1/3">papers homepage</a>:</p>

<p><img src="/assets/images/cbStillWorks.png" alt="" /></p>

<p>Now, the script has not yet been updated for the new journal (Noel, can you please upload the revision?), so you need
to edit the source right now and add <code class="language-plaintext highlighter-rouge">http://*.jcheminf.com/*</code> to the list of website the script acts on:</p>

<p><img src="/assets/images/cbStillWorks1.png" alt="" /></p>]]></content><author><name>Egon Willighagen</name></author><category term="cb" /><category term="cheminf" /><category term="cml" /><category term="userscript" /><category term="publishing" /><category term="rdf" /><category term="jcheminf" /><category term="justdoi:10.1186/1758-2946-1-2" /><category term="justdoi:10.1186/1758-2946-1-3" /><category term="doi:10.1186/1471-2105-8-487" /><summary type="html"><![CDATA[Besides Nature Chemistry , another journal was launched last week (see here and here): the Journal of Cheminformatics. First of all, congratulations to Chris and David for their efforts! While the journal only published one research paper yet, it already found its place on Chemical blogspace. I have two things I want to blog about: data rich publishing, and starting the scientific communication.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cbStillWorks.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cbStillWorks.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Autogenerating CML bindings for XMPP services with XMLBeans</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/03/14/autogenerating-cml-bindings-for-xmpp.html" rel="alternate" type="text/html" title="Autogenerating CML bindings for XMPP services with XMLBeans" /><published>2009-03-14T00:00:00+00:00</published><updated>2009-03-14T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/03/14/autogenerating-cml-bindings-for-xmpp</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/03/14/autogenerating-cml-bindings-for-xmpp.html"><![CDATA[<p>I blogged earlier about our efforts to create a better <a href="http://en.wikipedia.org/wiki/SOAP">SOAP</a>
service architecture, based on <a href="http://en.wikipedia.org/wiki/Jabber">XMPP</a>:</p>

<ul>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2009/01/21/details-behind-calling-xmpp-cloud.html">Details behind the “Calling XMPP cloud services from Taverna2” <i class="fa-solid fa-recycle fa-xs"></i></a></li>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2009/01/19/calling-xmpp-cloud-services-from.html">Calling XMPP cloud services from Taverna2 <i class="fa-solid fa-recycle fa-xs"></i></a></li>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2008/10/31/next-generation-asynchronous.html">Next generation asynchronous webservices <i class="fa-solid fa-recycle fa-xs"></i></a></li>
</ul>

<p>So, I set up XMPP services for QSAR descriptor calculation, 2D diagram and 3D geometry
calculations and a few more, using the <a href="http://cdk.sf.net/">CDK</a>.
<a href="http://en.wikipedia.org/wiki/Chemical_Markup_Language">Chemical Markup Language</a> has been my
primary choice for some 10 years now (see <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=1241">Peter’s blog</a>)
as it allows me to do things I cannot do in other formats.</p>

<p>Now, our XMPP services publish themselves what data types the allow as input and what they output
in return. They do this by publishing XML Schema to describe the input and output types. My CDK
services use CML, so they return the CML schema. Johannes’ <a href="http://xws4j.sourceforge.net/">xws4j</a>
implementation of the <a href="http://xmpp.org/extensions/xep-0244.html">IO-DATA</a>
specification has an add on that can build bindings to the schema on the fly. Now, CML comes with
a good <a href="http://www.xom.nu/">XOM</a>-based binding (called <a href="http://wwmm.ch.cam.ac.uk/maven2/cml/cmlxom/">CMLXOM</a>)
so this is not strictly necessary, but for less common schemata it is worthwhile: you can always
create bindings for brand new schemata, for older versions, for whatever. Services can even
create their own local schemata, and people will still be able to easily use them. This is to me
a big plus for this architecture.</p>

<p>Anyway, while CMLXOM exists, we wanted to show that the on-the-fly creation of bindings works,
even for large schemata, such as CML. However, one of the older flavours had an small error in a
regular expression in a data type CML defines. Johannes therefore asked me to test building
bindings for the CML schema version used in my services. He adviced me to use scomp for this,
which is a command line utility around the <a href="http://xmlbeans.apache.org/">XMLBeans</a>
library used for the binding generation.</p>

<p>As I am running <a href="http://www.ubuntu.com/">Ubuntu</a>, I preferred installing
<a href="http://packages.ubuntu.com/jaunty/libxmlbeans-java">the packaged version</a> instead of installing
the binary provided by XMLBeans. Now, after I did this, I noticed that this .deb did not install
the scomp utility, so I filled a <a href="https://bugs.launchpad.net/ubuntu/+source/xmlbeans/+bug/342349">wishlist bug report</a>.
Earlier this week I already encountered another bug, but this package being Java, I had a good
idea on how to fix the bug.</p>

<p>And so I implemented my own wishlist. I’m sure there is room for improvement, as my .deb
packaging skills are a bit rusty (a very long time ago I have been in the Debian New Maintainers
queue, but by the time they solved the long queue delays, I was too occupied with other things.
Yes, this was a long time ago already :). Anyway, Ubuntu’s <a href="http://launchpad.net/">LaunchPad</a>
has a nice feature, called the <a href="http://launchpad.net/ubuntu/+ppas">Personal Package Archives</a>.
This service will, after I have finished hacking on the packaging specs in the famous <em>debian/</em>
folder and tested the <em>.debs</em> build from it, will rebuild it and put the resulting package up for
download.</p>

<p>Conclusion: a perfect opportunity to finally gives this a try. The learning curve was
surprisingly shallow, and the result can be seen in <a href="https://launchpad.net/~egonw/+archive/ppa">my personal package archive</a>:</p>

<p><img src="/assets/images/ppa.png" alt="" /></p>

<p>Now, you can easily imagine that I will soon work on packaging stuff I did in the past too, such
as update <a href="http://packages.ubuntu.com/jaunty/libcdk-java">libcdk-java</a> and now that OpenJDK in
main can run <a href="http://www.jmol.org/">Jmol</a> reasonably, finally package Jmol for main. I just hope
I remember my <a href="http://alioth.debian.org/">Alioth</a> account, so that I can properly contribute to
the <a href="http://alioth.debian.org/projects/debichem/">debichem</a> project.</p>

<p>Getting back to running <em>scomp</em> on the CML scheme, it works with one minor problem:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span>scomp <span class="nt">-src</span> <span class="nb">.</span> <span class="nt">-d</span> <span class="nb">.</span>  cml.xsd
/home/egonw/tmp/cml/cml.xsd:10098:9: warning: p-props-correct.2.2: maxOccurs must be greater than or equal to 1.
Time to build schema <span class="nb">type </span>system: 1.792 seconds
Time to generate code: 3.297 seconds
Time to compile code: 9.658 seconds
</code></pre></div></div>

<p>The problem is reflected by line 10098 which goes like:</p>

<div class="language-xml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">&lt;xsd:sequence</span> <span class="na">minOccurs=</span><span class="s">"0"</span> <span class="na">maxOccurs=</span><span class="s">"0"</span><span class="nt">&gt;</span>
</code></pre></div></div>

<p>which can be traced down to line 23 in <a href="http://cml.svn.sf.net/viewvc/cml/schema2/trunk/elements/tableHeaderCell.xsd?revision=161&amp;view=markup">schema2/trunk/elements/tableHeaderCell.xsd</a>.
I filled a <a href="https://sourceforge.net/tracker2/?func=detail&amp;aid=2686810&amp;group_id=51361&amp;atid=463005">bug report about this</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cml" /><category term="java" /><category term="ubuntu" /><category term="xml" /><category term="xmpp" /><summary type="html"><![CDATA[I blogged earlier about our efforts to create a better SOAP service architecture, based on XMPP:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/ppa.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/ppa.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Editing and Validation of CML documents in Bioclipse</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/12/30/editing-and-validation-of-cml-documents.html" rel="alternate" type="text/html" title="Editing and Validation of CML documents in Bioclipse" /><published>2008-12-30T00:10:00+00:00</published><updated>2008-12-30T00:10:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/12/30/editing-and-validation-of-cml-documents</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/12/30/editing-and-validation-of-cml-documents.html"><![CDATA[<p>One advantage of using XML is that one can rely on good support in libraries for functionality. When
parsing XML, one does not have to take care of the syntax, and focus on the data and its semantics.
This comes at the expense of verbosity, though, but having the ability to express semantics explicitly
is a huge benefit for flexibility.</p>

<p>So, when <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/">Peter</a> and Henry put their first documents online about the Chemical Markup Language
(CML), I was thrilled, even though is actually was still SGML when I encountered it. The work predates the
<a href="http://www.w3.org/TR/1998/REC-xml-19980210">XML recommendation</a>. As I
<a href="https://chem-bla-ics.linkedchemistry.info/2008/10/02/jchempaint-history-cml-patches-in-1999.html">recently blogged <i class="fa-solid fa-recycle fa-xs"></i></a>, in ‘99
I wrote patches for Jmol and JChemPaint to support CML, which were published as preprint in the
<a href="http://www.sciencedirect.com/preprintarchive">Chemical Preprint Server</a> in a paper in 2000 in the
<a href="http://hackberry.trinity.edu/IJC/">Internet Journal of Chemistry</a>. Neither of the two has survived.</p>

<p>Anyway, the <a href="http://cdk.sf.net/">Chemistry Development Kit</a> makes heavy use of CML, and 
<a href="http://www.bioclipse.net/">Bioclipse</a> supports it too. Now, Bioclipse is based on the <a href="http://www.eclipse.org/">Eclipse</a>
<a href="http://wiki.eclipse.org/index.php/Rich_Client_Platform">Rich Client Platform</a> architecture, for which
there exist quite a few XML tools in the <a href="http://www.eclipse.org/webtools/">Web Tools Platform</a> (WTP).
Among these, a validation, content assisting XML editor. This means, I get red markings when I make my
XML document not-well-formed or invalid. Just a quick recap: well-formedness means that the XML document
has a proper syntax: one root node, properly closed tags, quotes around attribute values, etc. Validness,
however, means that the document is well-formed, but also hierarchically organized according to some specification.</p>

<p>Enter CML. CML is such a specification, first with DTDs, but after the introduction of XML Namespaces with
XML Schema (see <a href="http://cmlexplained.blogspot.com/2007/06/there-can-be-only-one-namespace.html">There can be only one (namespace)</a>).
The WTP can use this XML Schema for validation, and this is of great help learning the CML language.
Pressing Ctrl-space in Bioclipse will now show you what allowed content can be added at the current character
position.</p>

<p>Yes, Bioclipse can do this now (in SVN, at least). This has been on my wishlist for at least two years now, but
never really found the right information. Now, three days ago <a href="http://intellectualcramps.blogspot.com/">David</a>
wrote about <a href="http://intellectualcramps.blogspot.com/2008/12/end-of-year-cramps.html">End of Year Cramps</a>
in which he describes some of his work on the WTP for autocomplete for XPath queries. He <em>see[s] a brighter
future for XML at eclipse over the next year. I hope that those in the eclipse and XML community will help
to continue to improve the basic support, so that first class commercial quality applications that leverage
this support can continue to be built.</em></p>

<p>That was enough statement for me to <a href="http://intellectualcramps.blogspot.com/2008/12/end-of-year-cramps.html?showComment=1230451020000#c4332753586396921531">ask in the comments</a>
on how to make the WTP XML editor aware of the CML XML Schema. It already picked up XML Schema’s with
<code class="language-plaintext highlighter-rouge">xsi:schemaLocation</code>, but I needed something to worked without such statements in the XML document itself.
David explained that me that I could use the <a href="http://intellectualcramps.blogspot.com/2008/12/end-of-year-cramps.html?showComment=1230498780000#c4628316622126916885">org.eclipse.wst.xml.catalog extension</a>.
This was really easy, and <a href="http://bioclipse.svn.sourceforge.net/viewvc/bioclipse/bioclipse2/trunk/plugins/net.bioclipse.cml/plugin.xml?r1=8101&amp;r2=8100&amp;pathrev=8101">commited to Bioclipse SVN</a> as:</p>

<div class="language-xml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">&lt;extension</span>
  <span class="na">point=</span><span class="s">"org.eclipse.wst.xml.core.catalogContributions"</span><span class="nt">&gt;</span>
  <span class="nt">&lt;catalogContribution&gt;</span>
    <span class="nt">&lt;uri</span> <span class="na">name=</span><span class="s">"http://www.xml-cml.org/schema"</span>
          <span class="na">uri=</span><span class="s">"schema24/schema.xsd"</span><span class="nt">/&gt;</span>
  <span class="nt">&lt;/catalogContribution&gt;</span>
<span class="nt">&lt;/extension&gt;</span>
</code></pre></div></div>

<p>However, that does not make the WTP XML editor available in the Bioclipse application yet. Not ever in
the “Open With”… So, I set up a <a href="http://bioclipse.svn.sourceforge.net/viewvc/bioclipse/bioclipse2/trunk/features/net.bioclipse.cml_feature/">CML Feature</a>.
After a follow up question, it turned out that the CML content type of Bioclipse was already a sub type of the
XML type (see ):</p>

<div class="language-xml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">&lt;extension</span>
  <span class="na">point=</span><span class="s">"org.eclipse.core.runtime.contentTypes"</span><span class="nt">&gt;</span>
  <span class="nt">&lt;content-type</span>
    <span class="na">base-type=</span><span class="s">"org.eclipse.core.runtime.xml"</span>
    <span class="na">id=</span><span class="s">"net.bioclipse.contenttypes.cml"</span>
    <span class="na">name=</span><span class="s">"Chemical Markup Language (CML)"</span>
    <span class="na">file-extensions=</span><span class="s">"cml,xml"</span>
    <span class="na">priority=</span><span class="s">"normal"</span><span class="nt">&gt;</span>
  <span class="nt">&lt;/content-type&gt;</span>
<span class="nt">&lt;/extension&gt;</span>
</code></pre></div></div>

<p>So, the only remaining problem was to actually get the WTP XML editor as part of the Bioclipse application.
The new CML Feature takes care of that (I hope the export and building the update site work too, but
that’s yet untested), by important the relevant plugins and features. Last night, however, I ended up with
one stacktrace which gave me little clue on which plugin I was still missing.</p>

<p>Therefore, I headed to #eclipse and actually met David of the blog that started this again. He asked
<a href="http://nitind.blogspot.com/">nitind</a> to think about it too, and they helped me pin down the issue.
This relevant bit of the stacktrace turned out to be:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Caused by: java.lang.IllegalStateException
 at org.eclipse.core.runtime.Platform.getPluginRegistry(Platform.java:774)
 at org.eclipse.wst.common.componentcore.internal.impl.WTPResourceFactoryRegistry$ResourceFactoryRegistryReader.(WTPResourceFactoryRegistry.java:275)
 at org.eclipse.wst.common.componentcore.internal.impl.WTPResourceFactoryRegistry.(WTPResourceFactoryRegistry.java:61)
 at org.eclipse.wst.common.componentcore.internal.impl.WTPResourceFactoryRegistry.(WTPResourceFactoryRegistry.java:55)
 ... 37 more
</code></pre></div></div>

<p>This refered to this bit of code of Eclipse’ Platform.java:</p>

<div class="language-java highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nc">Bundle</span> <span class="n">compatibility</span> <span class="o">=</span> <span class="nc">InternalPlatform</span><span class="o">.</span><span class="na">getDefault</span><span class="o">()</span>
  <span class="o">.</span><span class="na">getBundle</span><span class="o">(</span><span class="nc">CompatibilityHelper</span><span class="o">.</span><span class="na">PI_RUNTIME_COMPATIBILITY</span><span class="o">);</span>
  <span class="k">if</span> <span class="o">(</span><span class="n">compatibility</span> <span class="o">==</span> <span class="kc">null</span><span class="o">)</span>
    <span class="k">throw</span> <span class="k">new</span> <span class="nf">IllegalStateException</span><span class="o">();</span>
</code></pre></div></div>

<p>So, the plugin I turned to to have missing was <em>org.eclipse.core.runtime.compatibility</em>. Apparently,
some parts of the WTP that the XMLEditor is using, still uses Eclipse2.x technology.</p>

<p><img src="/assets/images/cmlValid.png" alt="" /></p>

<p>This screenshot shows the WTP XMLEditor in action in Bioclipse on a CML file. It shows the document
contents with the ‘Design’ tab, which also shows allowed content, as derived from the XML Schema for
CML. Also, note that the Outline and Properties view automatically come for free, which allows more
detail and navigation of the content.</p>

<p><img src="/assets/images/cmlContentAssisting.png" alt="" /></p>

<p>This screenshot shows the ‘Source’ tab for the same file, where I deliberately changed the value of
the @id attribute of the first atom. The value does not validate against the regular expression defined
in the CML schema for @id attribute values. It also shows the content assisting in action. At any
location in the CML file, I can hit Ctrl-Space, and the editor will show me which content I can add
at that location.</p>

<p>This makes Bioclipse a perfect tool to craft CML documents and learn the language.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cml" /><category term="bioclipse" /><category term="xml" /><category term="cdk" /><summary type="html"><![CDATA[One advantage of using XML is that one can rely on good support in libraries for functionality. When parsing XML, one does not have to take care of the syntax, and focus on the data and its semantics. This comes at the expense of verbosity, though, but having the ability to express semantics explicitly is a huge benefit for flexibility.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cmlValid.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cmlValid.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Open{Data|Source|Standards} is not enough: we need Open Projects</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough.html" rel="alternate" type="text/html" title="Open{Data|Source|Standards} is not enough: we need Open Projects" /><published>2008-11-07T00:00:00+00:00</published><updated>2008-11-07T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough.html"><![CDATA[<p>The <a href="http://blueobelisk.sourceforge.net/wiki/Main_Page">Blue Obelisk</a> mantra <a href="http://blueobelisk.sourceforge.net/wiki/ODOSOS">ODOSOS</a>,
Open Data, Open Source, Open Standards, is well known, and much cited too. <a href="http://usefulchem.blogspot.com/">Jean-Claude Bradley</a>
popularized the <a href="http://en.wikipedia.org/wiki/Open_Notebook_Science">Open Notebook Science</a> (ONS). This has always been nagging me a bit,
because the <a href="http://cdk.sf.net/">CDK</a>, <a href="http://www.jmol.org/">Jmol</a>, JChemPaint and other chemistry projects have done that for much
longer, though we did not use notebooks as much, so called it just an open source project. It really is no different, IMO, though
surely, there are differences.</p>

<p>Anyway, the key thing which ONS and CDK and Jmol share, is that they use an Open Notebook. Not every Open Source or Open Data project does.
Actually, many scientific Open Source are not open Projects! They are more like the Cathedral than the wished-for Bazaar (see
<a href="http://en.wikipedia.org/wiki/The_Cathedral_and_the_Bazaar">The Cathedral and the Bazaar</a>). So, Open Source (science) projects are certainly not ONS projects by default!</p>

<p>Now, the CDK actually is ONS, it is a Bazaar. The notebooks we use include:</p>

<ul>
  <li>open project via <a href="https://sourceforge.net/mail/?group_id=20024">mailing lists</a></li>
  <li>open methods/results via <a href="https://sourceforge.net/svn/?group_id=20024">subversion</a></li>
  <li>informal reporting via blogs (e.g. <a href="http://rguha.wordpress.com/">Rajarshi</a>, <a href="http://www.steinbeck-molecular.de/steinblog/">Christoph</a>, <a href="http://cdktaverna.wordpress.com/">Thomas</a>, mine)</li>
  <li>informal reporting via <a href="http://www.cdknews.org/">CDK News</a></li>
</ul>

<p>What more would you wish for? That’s not a rhetorical question. Remember that every reader of this blog is in
<a href="https://chem-bla-ics.linkedchemistry.info/2007/11/27/be-in-my-advisory-board-1-being-good.html">my advisory board <i class="fa-solid fa-recycle fa-xs"></i></a>!</p>

<p>Unfortunately, I do not create work at a workbench myself, so I do not produce new knowledge myself, other than extracted from existing
data. That’s really a shame, and I really do hope that Jean-Claude or <a href="http://blog.openwetware.org/scienceintheopen">Cameron</a> will send
me a box to measure solubilities (see <a href="http://usefulchem.blogspot.com/2008/10/rdf-triples-for-open-notebook-science.html">here</a>,
<a href="http://usefulchem.blogspot.com/2008/11/ons-solubility-web-query.html">here</a>, and
<a href="http://anybody.cephb.fr/perso/lindenb/tmp/jcbradley.rdf">here</a>,
<a href="http://rguha.wordpress.com/2008/11/06/solubility-queries-and-the-google-visualization-api/">here</a> for first data exploration),
even though I cannot participate in the <a href="http://usefulchem.blogspot.com/2008/11/submeta-open-notebook-science-awards.html">challenge</a>.
(hint, hint :)</p>

<h2 id="from-cathedral-to-bazaar-in-life-sciences">From Cathedral to Bazaar in Life Sciences</h2>

<p>One Cathedral we ran into with <a href="http://www.bioclipse.net/">Bioclipse</a> was <a href="http://www.biocatalogue.org/">BioCatalogue</a>,
which will serve as website where people can annotate and categorize (web) services. While the project has been around for a while, the
website was rather uninformative. Fortunately, the projects is going to open up, and be more Bazaar-like. For example, they
now started a <a href="http://www.biocatalogue.org/wiki">wiki</a> and a
<a href="http://listserv.manchester.ac.uk/cgi-bin/wa?SUBED1=biocatalogue-friends&amp;A=1">mailing list</a>. I hope these efforts will continue,
so that I can contribute from my point of view!</p>

<p>The <a href="http://embraceregistry.net/">EMBRACE Registry</a> is a project with similar goals and a rather nice outcome (which I learned about on
<a href="https://chem-bla-ics.linkedchemistry.info/2008/11/03/embrace-workshop-in-uppsala.html">Monday <i class="fa-solid fa-recycle fa-xs"></i></a>). It is actually anticipate to be replaced by or merge
with BioCatalogue. So, all data I entered, <a href="http://prints.cs.man.ac.uk:8081/category/tags/cheminformatics">cheminformatics workflows</a>
(look, <a href="https://chem-bla-ics.linkedchemistry.info/2008/10/18/chemoinformatics-p0wned-by.html">no ‘o’ <i class="fa-solid fa-recycle fa-xs"></i></a>), will later be available from BioCatalogue too.
That is already my first contribution to BioCatalogue. One enormously interesting feature of the Registry, is that is allows uploading of
code to test the service. This will mean the Registry will not only poll if the service is still online (by checking the WSDL file), it
will also test if the service behaves properly. Now, immediate thoughts are mashups with <a href="http://www.myexperiment.org/">MyExperiment</a>.
Each WSDL entry in the Registry points to MyExperiment workflows that use them, and the workflow page would indicate the status of all
used WDSL services. This integration was already anticipated long before I thought about it, as the involved Cathedrals were nicely
located in the same floor in Manchester.</p>

<p>Below is a screenshot from the EMBRACE Registry for the <a href="http://www.chemspider.com/">ChemSpider</a>
<a href="http://prints.cs.man.ac.uk:8081/service/massspecapi">WDSL entry</a> for <a href="http://www.myexperiment.org/workflows/97">a workspace</a>
I <a href="https://chem-bla-ics.linkedchemistry.info/2007/11/26/metabolomics-workflows-in-taverna.html">uploaded <i class="fa-solid fa-recycle fa-xs"></i></a> about a year ago to MyExperiment:</p>

<p><img src="/assets/images/registry.png" alt="" /></p>

<p>BTW, ChemSpider has an Advisory Board of which I am member, but it is also a classical (and intentional) Cathedral project. We do share common interests though, which makes us collaborate.</p>

<h2 id="why-important">Why Important?</h2>

<p>One recurrent theme in Open Source is <a href="http://en.wikipedia.org/wiki/Given_enough_eyeballs">given enough eyeballs, all bugs are shallow</a>.
This surely applies to science as well. The difference between the two is that in current science the eyes only inspect with a delay of at
least 6 months. Current practice is that research is finished (delay), and when decided publishable written up a paper (delay, and loosing
valuable information in the process, as you can read in my blog all the time), and published (even more delay). ONS changes that, and so do
Bazaar-like open source projects, such as the CDK, Jmol and Bioclipse. They bugs are present, whether we like it or not, not just in source
code, but in science too. Theories get overthrown, but why should we like the long delays current scientific good practice? Hate it! Work
around it. Use the Bazaar. Use ONS!</p>

<p>Now, ONS actually needs Open Source, allowing them to deal effectively with the data they produce; to allow extraction of new scientific
knowledge from the measurements. If Rajarshi and Pierre would not have made their efforts, other could not easily join in, leading to
those much hated delays. Bugs should be shallow, and openness allows us to make those bugs visible. We can prove that there is a bug,
without having to reproduce data ourselves, leading to those nasty delays again. Just copy the data, compare it to your own, do your
analysis.</p>

<p>One recent project in open source chemistry dealing with making bugs visible, is the web page set up by Andreas Tille for the
<a href="http://alioth.debian.org/projects/debichem">DebiChem project</a>. His page <a href="http://cdd.alioth.debian.org/debichem/bugs/">summarizes the bugs</a>
listed for the chemistry in Debian (which includes the Blue Obelisk projects <a href="http://packages.debian.org/lenny/avogadro">Avogadro</a>,
<a href="http://packages.debian.org/lenny/bodr">BODR</a>, <a href="http://packages.debian.org/lenny/libcdk-java">CDK</a>,
<a href="http://packages.debian.org/lenny/chemical-mime-data">Chemical MIME Data</a>,
<a href="http://packages.debian.org/lenny/kalzium">Kalzium</a> and <a href="http://packages.debian.org/lenny/openbabel">OpenBabel</a>):</p>

<p><img src="/assets/images/debichem.png" alt="" /></p>

<p>This data analysis helps the projects being analyzed.</p>

<h2 id="packaging">Packaging</h2>

<p>This brings me to a last topic, for this blog: packaging using Open Standards. In order to allow those eyeballs to spot bugs, it is of the
utmost importance to package your results in Open Standards, and not just one, but likely many. For Open Source projects this ultimately
means Distribution Packages (deb or rpm). If that goal has been achieved, you know your results can be read by anyone. Software should be
installable (make, ant, cmake, etc), and Data should be readable (no PDF, but RDF, XML, JSON, or whatever standard). Preferably not Excel,
as this is too free format (as Rajarshi also <a href="http://rguha.wordpress.com/2008/11/06/solubility-queries-and-the-google-visualization-api/">indicated</a>),
but with some added conventions it may do well. Blue Obelisk project are generally doing well in terms of packaging.</p>

<p>For the CDK, which already is reasonably well packaged, I am currently working on <a href="http://cdk.svn.sourceforge.net/viewvc/cdk/cdk-eclipse/trunk/">Eclipse</a>
and <a href="http://cdk.svn.sourceforge.net/viewvc/cdk/cdk-pom/trunk/">Maven2</a> packages. The former is already being used by Bioclipse, while the
second aims at <a href="https://sourceforge.net/projects/cml">Jumbo</a> (which has just seen a
<a href="https://sourceforge.net/project/showfiles.php?group_id=51361">new release</a>. <a href="http://wwmm.ch.cam.ac.uk/blogs/downing/">Jim</a>,
I’m happy to see the CMLDOM/Jumbo split!), <a href="http://www.cdk-taverna.de/">CDK-Taverna</a>, and possibly a third (Paula, what for do you plan
to use it?). The POM export is not fully working yet, but with four research sites involved in this Open Project, I’m sure we’ll work
it out.</p>

<p>The bottom line is, scientific progress would benefit so much from a Bazaar approach. And the key thing is not collaboration; that’s
something you can do in a Cathedral-like fashion too. No, the key thing is to be Open and allow anyone, even your worst nightmare, to
comment on what you do. Let him prove you wrong, openly, that is.</p>

<p>OK, there it is. My open notebook entry for this week. Now you know what I have been up to this week.</p>]]></content><author><name>Egon Willighagen</name></author><category term="odosos" /><category term="chemspider" /><category term="workflow" /><category term="cdk" /><category term="bioclipse" /><category term="cml" /><category term="debian" /><category term="eclipse" /><category term="rdf" /><category term="jmol" /><category term="blue-obelisk" /><summary type="html"><![CDATA[The Blue Obelisk mantra ODOSOS, Open Data, Open Source, Open Standards, is well known, and much cited too. Jean-Claude Bradley popularized the Open Notebook Science (ONS). This has always been nagging me a bit, because the CDK, Jmol, JChemPaint and other chemistry projects have done that for much longer, though we did not use notebooks as much, so called it just an open source project. It really is no different, IMO, though surely, there are differences.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/registry.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/registry.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Bioclipse2 Scripting #1: from SMILES to a UFF optimized structure in Jmol</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/10/25/bioclipse2-scripting-1-from-smiles-to.html" rel="alternate" type="text/html" title="Bioclipse2 Scripting #1: from SMILES to a UFF optimized structure in Jmol" /><published>2008-10-25T00:00:00+00:00</published><updated>2008-10-25T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/10/25/bioclipse2-scripting-1-from-smiles-to</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/10/25/bioclipse2-scripting-1-from-smiles-to.html"><![CDATA[<p>After some difficulties this week with making an export of <a href="http://cdk.sf.net/">CDK</a> plugins in the
<a href="http://www.bioclipse.net/">Bioclipse2</a> <em>Cheminformatics feature</em> of with the <a href="http://cdk.svn.sourceforge.net/viewvc/cdk/cdk-eclipse/trunk/">cdk-eclipse</a>
software, I got the following cute Bioclipse2 script up and running:</p>

<div class="language-javascript highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nx">dimethylether</span> <span class="o">=</span> <span class="nx">cdk</span><span class="p">.</span><span class="nf">fromSMILES</span><span class="p">(</span> <span class="dl">"</span><span class="s2">COC</span><span class="dl">"</span> <span class="p">);</span>
<span class="nx">cdk</span><span class="p">.</span><span class="nf">addExplicitHydrogens</span><span class="p">(</span> <span class="nx">dimethylether</span> <span class="p">);</span>
<span class="nx">cdk</span><span class="p">.</span><span class="nf">generate3dCoordinates</span><span class="p">(</span> <span class="nx">dimethylether</span> <span class="p">);</span>

<span class="c1">// save as CML</span>
<span class="nx">cdk</span><span class="p">.</span><span class="nf">saveCML</span><span class="p">(</span> <span class="nx">dimethylether</span><span class="p">,</span> <span class="dl">"</span><span class="s2">/Virtual/dimethylether.cml</span><span class="dl">"</span> <span class="p">);</span>
<span class="nx">ui</span><span class="p">.</span><span class="nf">open</span><span class="p">(</span> <span class="dl">"</span><span class="s2">/Virtual/dimethylether.cml</span><span class="dl">"</span> <span class="p">);</span> <span class="c1">// this should open a JmolEditor</span>

<span class="nx">jmol</span><span class="p">.</span><span class="nf">minimize</span><span class="p">();</span>
</code></pre></div></div>

<p>You can see four of my favorite cheminformatics tools integrated: CDK is used to convert a SMILES into connection table with add explicit
hydrogens, and to create initial 3D coordinates (with the code from Christian Hoppe, and thanx to Stefan for fixing that code in the
CDK 1.1.x branch!). Then, <a href="http://cml.sourceforge.net/">CMLDOM</a> is used to create and save a CML document, which is then opened into a
<a href="http://www.jmol.org/">Jmol</a> editor in Bioclipse.</p>

<p>A variation of this script is visible in the following screenshot:</p>

<p><img src="/assets/images/mashupCmldomJmolCDK.png" alt="" /></p>

<p>This and other Bioclipse2 scripts I will post in <a href="http://gist.github.com/">Gist</a>, a sort of <a href="http://en.wikipedia.org/wiki/Pastebin">pastebin</a>
supporting version history, and I’ll tag them with <em>bioclipse gist</em> on <a href="http://delicious.com/egonw/">delicious</a>, so that you can always
browse them, comment on them, or add your own gists at
<a href="http://delicious.com/tag/bioclipse+gist">http://delicious.com/tag/bioclipse+gist</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="bioclipse" /><category term="cml" /><category term="cdk" /><category term="jmol" /><category term="eclipse" /><category term="github" /><category term="cheminf" /><summary type="html"><![CDATA[After some difficulties this week with making an export of CDK plugins in the Bioclipse2 Cheminformatics feature of with the cdk-eclipse software, I got the following cute Bioclipse2 script up and running:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/mashupCmldomJmolCDK.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/mashupCmldomJmolCDK.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">JChemPaint history: CML patches in 1999</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/10/02/jchempaint-history-cml-patches-in-1999.html" rel="alternate" type="text/html" title="JChemPaint history: CML patches in 1999" /><published>2008-10-02T00:00:00+00:00</published><updated>2008-10-02T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/10/02/jchempaint-history-cml-patches-in-1999</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/10/02/jchempaint-history-cml-patches-in-1999.html"><![CDATA[<p>There was some talk about the history of chemoinformatics toolkits by
<a href="http://baoilleach.blogspot.com/2008/09/overview-of-cheminformatics-toolkits.html">Noel</a> and
<a href="http://www.dalkescientific.com/writings/diary/archive/2008/09/20/euroqsar.html">Andrew</a>, which made
me wonder on the exact history of <a href="http://www.jmol.org/">Jmol</a> and
<a href="http://sf.net/project/jchempaint">JChemPaint</a>. Below is the email
<a href="http://www.steinbeck-molecular.de/steinblog/">Christoph</a> dug up from his archives:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>X-Mozilla-Status: 1011
X-Mozilla-Status2: 00000000
Message-ID: &lt;372ECD5E.53A49584@ice.mpg.de&gt;
Date: Tue, 04 May 1999 12:35:10 +0200
From: Christoph Steinbeck
Reply-To: steinbeck@ice.mpg.de
Organization: Max-Planck-Institute of Chemical Ecology
X-Mailer: Mozilla 4.51 [en] (WinNT; I)
X-Accept-Language: en
MIME-Version: 1.0
To: Egon Willighagen
Subject: Re: Participating in JChemPaint
References: &lt;000701be9613$34cf52e0$8e74ae83@catv6142.extern.kun.nl&gt;
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit

&gt; Egon Willighagen wrote:
&gt;
&gt; Dear Christoph Steinbeck,
&gt;
&gt; Yesterday I visited your site on JChemPaint. I like to contribute some
&gt; of my expertise on
&gt; Java and CML (1).
&gt;
&gt; CML is a markup language that is able to contain chemical information.
&gt; It can contain for example physical properties, for which I use CML in
&gt; my Dictionary on Organic Chemistry (2).
&gt; But is also might contain spectra, bibliographic references etc. And
&gt; of course 2D and 3D
&gt; structural information.
&gt;
&gt; Therefore I propose to write both CML-input and -output procedures for
&gt; the JChemPaint project.
&gt;
&gt; I hope to hear from you soon.
&gt;
&gt; Yours sincerely,
&gt;
&gt; Egon Willighagen
&gt;
&gt; 1. http://www.xml-cml.org/
&gt; 2. http://www.sci.kun.nl/sigma/Chemisch/Woordenboek/

Dear Egon,

thanks very much for your mail and your offer to write CML-input and
output routines for JChemPaint.
That really sounds great to me and I will give you access to our CVS
tree as soon as we have discussed the details.

Cheers,

Chris

--C. S.
Dr. Christoph Steinbeck (http://www.ice.mpg.de/~stein)
MPI of Chemical Ecology, Tatzendpromenade 1a, 07745 Jena, Germany
Tel: +49(0)3641 643644 - MoPho: +49(0)177 8236510 - Fax: +49(0)3641
643665

What is man but that lofty spirit - that sense of enterprise.
.. Kirk, "I, Mudd," stardate 4513.3..
</code></pre></div></div>

<p>Now, my email must have been triggered by the <a href="http://freshmeat.net/projects/jchempaint/">announcement of JChemPaint on FreshMeat.net</a>,
which is the oldest public record of JChemPaint I have found so far:</p>

<p><img src="/assets/images/fmJChemPaint.png" alt="" /></p>]]></content><author><name>Egon Willighagen</name></author><category term="jmol" /><category term="jchempaint" /><category term="cml" /><summary type="html"><![CDATA[There was some talk about the history of chemoinformatics toolkits by Noel and Andrew, which made me wonder on the exact history of Jmol and JChemPaint. Below is the email Christoph dug up from his archives:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/fmJChemPaint.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/fmJChemPaint.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Creating CMLReact from UsefulChem Ugi Reactions</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/08/31/creating-cmlreact-from-usefulchem-ugi.html" rel="alternate" type="text/html" title="Creating CMLReact from UsefulChem Ugi Reactions" /><published>2008-08-31T00:20:00+00:00</published><updated>2008-08-31T00:20:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/08/31/creating-cmlreact-from-usefulchem-ugi</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/08/31/creating-cmlreact-from-usefulchem-ugi.html"><![CDATA[<p><a href="http://blog.openwetware.org/scienceintheopen/">Cameron</a>, <a href="http://usefulchem.blogspot.com/">Jean-Claude</a> and I were invited to
<a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/">Peter</a>’s place in Cambridge, where we are now hacking on CMLReact for the
<a href="http://usefulchem.wikispaces.com/exp023">Ugi reactions</a> Jean-Claude has been working on. I just finished a script that uses the
CDK and Sam’s interface to the <a href="http://cheminfo.informatics.indiana.edu/~rguha/code/java/nightly/api/org/openscience/cdk/inchi/package-frame.html">InChI library</a>
to convert a list of four reactants and one Ugi product into CMLReact (doi:<a href="http://dx.doi.org/10.1021/ci0502698">10.1021/ci0502698</a>).
The full <a href="https://en.wikipedia.org/wiki/BeanShell">BeanShell</a> script looks like:</p>

<pre><code class="language-beanshell">#!/usr/bin/bsh

import java.io.File;
import java.io.FileReader;
import java.io.BufferedReader;

import org.openscience.cdk.*;
import org.openscience.cdk.exception.*;
import org.openscience.cdk.inchi.*;
import org.openscience.cdk.interfaces.*;
import org.openscience.cdk.io.CMLWriter;

import org.openscience.cdk.libio.cml.Convertor;
import org.xmlcml.cml.element.CMLReaction;

import net.sf.jniinchi.INCHI_RET;

InChIGeneratorFactory factory = new InChIGeneratorFactory();
// Get InChIToStructure

File file = new File("inchi.ugi.txt"); // five inchis expected, last being the product
BufferedReader reader = new BufferedReader(new FileReader(file));

String first = reader.readLine();
String second = reader.readLine();
String third = reader.readLine();
String fourth = reader.readLine();
String product = reader.readLine();

System.out.println("First: " + first);
IMolecule firstAC;
{
  InChIToStructure intostruct = factory.getInChIToStructure(first, DefaultChemObjectBuilder.getInstance());

  INCHI_RET ret = intostruct.getReturnStatus();
  if (ret == INCHI_RET.WARNING) {
    // Structure generated, but with warning message
    System.out.println("InChI warning: " + intostruct.getMessage());
  } else if (ret != INCHI_RET.OKAY) {
    // Structure generation failed
    throw new CDKException("Structure generation failed failed: " + ret.toString()
      + " [" + intostruct.getMessage() + "]");
  }

  firstAC = new Molecule(intostruct.getAtomContainer());
}

System.out.println("Second: " + second);
IMolecule secondAC;
{
  InChIToStructure intostruct = factory.getInChIToStructure(second, DefaultChemObjectBuilder.getInstance());

  INCHI_RET ret = intostruct.getReturnStatus();
  if (ret == INCHI_RET.WARNING) {
    // Structure generated, but with warning message
    System.out.println("InChI warning: " + intostruct.getMessage());
  } else if (ret != INCHI_RET.OKAY) {
    // Structure generation failed
    throw new CDKException("Structure generation failed failed: " + ret.toString()
      + " [" + intostruct.getMessage() + "]");
  }

  secondAC = new Molecule(intostruct.getAtomContainer());
}

System.out.println("Third: " + third);
IMolecule thirdAC;
{
  InChIToStructure intostruct = factory.getInChIToStructure(third, DefaultChemObjectBuilder.getInstance());

  INCHI_RET ret = intostruct.getReturnStatus();
  if (ret == INCHI_RET.WARNING) {
    // Structure generated, but with warning message
    System.out.println("InChI warning: " + intostruct.getMessage());
  } else if (ret != INCHI_RET.OKAY) {
    // Structure generation failed
    throw new CDKException("Structure generation failed failed: " + ret.toString()
      + " [" + intostruct.getMessage() + "]");
  }

  thirdAC = new Molecule(intostruct.getAtomContainer());
}

System.out.println("Fourth: " + fourth);
IMolecule fourthAC;
{
  InChIToStructure intostruct = factory.getInChIToStructure(fourth, DefaultChemObjectBuilder.getInstance());

  INCHI_RET ret = intostruct.getReturnStatus();
  if (ret == INCHI_RET.WARNING) {
    // Structure generated, but with warning message
    System.out.println("InChI warning: " + intostruct.getMessage());
  } else if (ret != INCHI_RET.OKAY) {
    // Structure generation failed
    throw new CDKException("Structure generation failed failed: " + ret.toString()
      + " [" + intostruct.getMessage() + "]");
  }

  fourthAC = new Molecule(intostruct.getAtomContainer());
}

System.out.println("Product: " + product);
IMolecule productAC;
{
  InChIToStructure intostruct = factory.getInChIToStructure(product, DefaultChemObjectBuilder.getInstance());

  INCHI_RET ret = intostruct.getReturnStatus();
  if (ret == INCHI_RET.WARNING) {
    // Structure generated, but with warning message
    System.out.println("InChI warning: " + intostruct.getMessage());
  } else if (ret != INCHI_RET.OKAY) {
    // Structure generation failed
    throw new CDKException("Structure generation failed failed: " + ret.toString()
      + " [" + intostruct.getMessage() + "]");
  }

  productAC = new Molecule(intostruct.getAtomContainer());
}

IReaction ugiReaction = new Reaction();
ugiReaction.addReactant(firstAC);
ugiReaction.addReactant(secondAC);
ugiReaction.addReactant(thirdAC);
ugiReaction.addReactant(fourthAC);
ugiReaction.addProduct(productAC);

StringWriter stringWriter = new StringWriter();
CMLWriter cmlWriter = new CMLWriter(stringWriter);

cmlWriter.write(ugiReaction);
cmlWriter.close();
System.out.println(stringWriter.toString());
</code></pre>

<p>My apologies for the code duplication, but never tried inline functions in BeanShell yet… You can
monitor the efforts at <a href="http://docs.google.com/Doc?id=dq5m5bs_12hb8d2wcw">Google Docs</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cml" /><category term="cdk" /><category term="inchi" /><category term="justdoi:10.1021/ci0502698" /><category term="usefulchem" /><summary type="html"><![CDATA[Cameron, Jean-Claude and I were invited to Peter’s place in Cambridge, where we are now hacking on CMLReact for the Ugi reactions Jean-Claude has been working on. I just finished a script that uses the CDK and Sam’s interface to the InChI library to convert a list of four reactants and one Ugi product into CMLReact (doi:10.1021/ci0502698). The full BeanShell script looks like:]]></summary></entry></feed>