<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemspider.xml" rel="self" type="application/atom+xml" /><link href="https://chem-bla-ics.linkedchemistry.info/" rel="alternate" type="text/html" /><updated>2026-08-14T05:35:03+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemspider.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">Compound (class) identifiers in Wikidata</title><link href="https://chem-bla-ics.linkedchemistry.info/2018/08/18/compound-class-identifiers-in-wikidata.html" rel="alternate" type="text/html" title="Compound (class) identifiers in Wikidata" /><published>2018-08-18T00:00:00+00:00</published><updated>2018-08-18T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2018/08/18/compound-class-identifiers-in-wikidata</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2018/08/18/compound-class-identifiers-in-wikidata.html"><![CDATA[<p><span style="width: 40%; display: block; margin-left: auto; margin-right: auto; float: right">
<img src="/assets/images/extid-wikidata-histogram.png" /> <br />
<a href="https://edu.nl/h6kg3">Bar chart</a> showing the number of compounds with a particular chemical identifier.
</span>
I think <a href="http://wikidata.org/">Wikidata</a> is a groundbreaking project, which will have a major impact on science. One of the
reasons is the open license (CCZero), the very basic approach (<a href="http://wikiba.se/">Wikibase</a>), and the superb community around
it. For example, setting up your own Wikibase including a cool SPARQL endpoint, is
<a href="https://github.com/wmde/wikibase-docker">easily done with Docker</a>.</p>

<p>Wikidata has many sub projects, such as <a href="http://wikicite.org/">WikiCite</a>, which captures the collective of primary literature.
Another one is the <a href="https://www.wikidata.org/wiki/Wikidata:WikiProject_Chemistry">WikiProject Chemistry</a>. The two nicely match
up, I think, making a public database linking chemicals to literature (tho, very much needs to be done here), see my recent
ICCS 2018 poster (doi:<a href="https://doi.org/10.6084/m9.figshare.6356027.v1">10.6084/m9.figshare.6356027.v1</a>, paper pending).</p>

<p>But Wikidata is also a great resource for identifier mappings between chemical databases, something we need for
<a href="https://chem-bla-ics.blogspot.com/2017/11/new-paper-wikipathways-multifaceted.html">our metabolism pathway research</a>.
The mapping, as you may know, are <a href="https://chem-bla-ics.blogspot.com/2016/09/metabolite-identifier-mapping-databases.html">used in the latter</a>
via <a href="https://www.bridgedb.org/">BridgeDb</a> and we have been using Wikidata as one of three sources for some time now (the others being
<a href="http://www.hmdb.ca/">HMDB</a> and <a href="https://www.ebi.ac.uk/chebi/">ChEBI</a>). WikiProject Chemistry has a related
<a href="https://www.wikidata.org/wiki/Wikidata:WikiProject_Chemistry/ChemID">ChemID</a> effort, and while the wiki page does not show
much recent activity, there is actually a lot of ongoing effort (see <a href="https://edu.nl/h6kg3">plot</a>).
And I’ve been <a href="https://chem-bla-ics.blogspot.com/2018/07/lipid-map-identifiers-and.html">adding my bits</a>.</p>

<h2 id="limitations-of-the-links">Limitations of the links</h2>
<p>But not each identifier in Wikidata has the same meaning. While they are all classified as ‘external-id’, the actual link may
have different meaning. This, of course, is the essence of scientific lenses, see <a href="https://chem-bla-ics.blogspot.com/2013/05/linking-wikipathways-to-binding.html">this post</a>
and the papers cited therein. One reason here is the difference in what entries in the various databases mean.</p>

<p>Wikidata has an extensive model, defined by the aforementioned WikiProject Chemistry. For example, it has different concepts
for chemical compounds (in fact, the hierarchy is pretty rich) and compound classes. And these are differently modeled. Furthermore,
it has a model that formalizes that things with a different InChI are different, but even allows things with the same InChI to be
different, if need arises. It tries to accurately and precisely capture the certainty and uncertainty of the chemistry. As such,
it is a powerful system to handle identifier mappings, because databases are not clear, and chemistry and biological in data is
even less: we measure experimentally a characterization of chemicals, but what we put in databases and give names, are specific
models (often chemical graphs).</p>

<p>That model differs from what other (chemical) databases use, or seem to use, because not always do databases indicate what they
actually have in a record. But I think this is a fair guess.</p>

<h2 id="chebi">ChEBI</h2>
<p>ChEBI (and the matching <a href="https://www.wikidata.org/wiki/Property:P683">ChEBI ID</a>) has entries for chemical classes (e.g.
<a href="https://www.ebi.ac.uk/chebi/searchId.do?chebiId=CHEBI:35366">fatty acid</a>) and specific compounds (e.g.
<a href="https://www.ebi.ac.uk/chebi/searchId.do?chebiId=30089">acetate</a>).</p>

<h2 id="pubchem-chemspider-unichem">PubChem, ChemSpider, UniChem</h2>
<p>These three resources use the InChI as central asset. While they do not really have the concept of compound classes so much
(though increasingly they have classifications), they do have entries where stereochemistry is undefined or unknown. Each
one has their own way to link to other databases themselves, which normally includes tons of structure normalization (see
e.g. doi:<a href="https://doi.org/10.1186/s13321-018-0293-8">10.1186/s13321-018-0293-8</a> and
doi:<a href="https://doi.org/10.1186/s13321-015-0072-8">10.1186/s13321-015-0072-8</a>).</p>

<h2 id="hmdb">HMDB</h2>
<p>HMDB (and the matching <a href="https://www.wikidata.org/wiki/Property:P2057">P2057</a>) has a biological perspective; the entries
reflect the biology of a chemical. Therefore, for most compounds, they focus on the neutral forms of compounds. This makes
linking to/from other databases where the compound is not neutral chemically less precise.</p>

<h2 id="cas-registry-numbers">CAS registry numbers</h2>
<p>CAS (and the matching <a href="https://www.wikidata.org/wiki/Property:P231">P231</a>) is pretty unique itself, and has identifiers
for substances (see <a href="https://www.wikidata.org/wiki/Q79529">Q79529</a>), much more than chemical compounds, and comes with a
own set of unique features. For example, solutions of some compound, by design, have the same identifier. Previously,
formaldehyde and formalin had different Wikipedia/Wikidata pages, both with the same CAS registry number.</p>

<h2 id="limitations-of-the-links-2">Limitations of the links #2</h2>
<p>Now, returning to our starting point: limitations in linking databases. If we want FAIR mappings, we need to be as precise
as possible. Of course, that may mean we need more steps, but we can always simplify at will, but we never can have a
computer make the links more complex (well, not without making assumptions, etc).</p>

<p>And that is why Wikidata is so suitable to link all these chemical databases: it can distinguish differences when needed,
and make that explicit. It make mappings between the databases more <a href="https://www.nature.com/articles/sdata201618">FAIR</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="wikidata" /><category term="scholia" /><category term="chemistry" /><category term="bridgedb" /><category term="cas" /><category term="chebi" /><category term="chemspider" /><category term="fair" /><category term="hmdb" /><category term="pubchem" /><category term="rdf" /><category term="wikicite" /><category term="justdoi:10.6084/m9.figshare.6356027.v1" /><category term="justdoi:10.1186/s13321-018-0293-8" /><category term="justdoi:10.1186/s13321-015-0072-8" /><category term="justdoi:10.1038/sdata.2016.18" /><summary type="html"><![CDATA[Bar chart showing the number of compounds with a particular chemical identifier. I think Wikidata is a groundbreaking project, which will have a major impact on science. One of the reasons is the open license (CCZero), the very basic approach (Wikibase), and the superb community around it. For example, setting up your own Wikibase including a cool SPARQL endpoint, is easily done with Docker.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/extid-wikidata-histogram.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/extid-wikidata-histogram.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">New Paper: “The ChEMBL database as linked open data”</title><link href="https://chem-bla-ics.linkedchemistry.info/2013/05/09/new-paper-chembl-database-as-linked.html" rel="alternate" type="text/html" title="New Paper: “The ChEMBL database as linked open data”" /><published>2013-05-09T00:00:00+00:00</published><updated>2013-05-09T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2013/05/09/new-paper-chembl-database-as-linked</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2013/05/09/new-paper-chembl-database-as-linked.html"><![CDATA[<script src="https://d1bxh8uas1mnw7.cloudfront.net/assets/embed.js" type="text/javascript"></script>

<div class="altmetric-embed" data-badge-details="right" data-badge-type="donut" data-doi="10.1186/1758-2946-5-23" style="float: right;"></div>

<p><strong>Update</strong>: Mark wrote up a <a href="http://chembl.blogspot.co.uk/2013/05/chembl-chembl-rdf.html">blog post</a> on the RDF that the ChEMBL team itself.</p>

<p>Yesterday, the paper “The ChEMBL database as linked open data” (doi:<a href="https://doi.org/10.1186/1758-2946-5-23">10.1186/1758-2946-5-23</a>) by
Andra Waagmeester (<a href="https://twitter.com/andrawaag">@andrawaag</a>), Ola Spjuth (<a href="https://twitter.com/ola_spjuth">@ola_spjuth</a>), Peter Ansell
(<a href="http://twitter.com/p_ansell">@p_ansell</a>), Antony Williams (<a href="https://twitter.com/chemconnector">@chemconnector</a>), Valery Tkachenko,
Janna Hastings, Bin Chen (<a href="http://twitter.com/binchenindiana">@binchenindiana</a>), David J Wild (<a href="http://twitter.com/davidjohnwild">@davidjohnwild</a>),
and me appeared in the OA <a href="http://en.wikipedia.org/wiki/Journal_of_Cheminformatics">JChemInf</a> journal.</p>

<p>I am also indebted to the <a href="https://www.ebi.ac.uk/chembl/">ChEMBL</a> team (<a href="http://twitter.com/chembl">@chembl</a>) for both providing such
valuable data under a liberal Open Access license and their critical reading of the manuscript! <strong>Additionally, I would like to stress
that the ChEMBL team will create their own RDF version of ChEMBL and that this paper is not describing the version they will release.</strong></p>

<p>BTW, the <a href="https://github.com/egonw/chembl-rdf-paper/">source of the paper</a> is available from GitHub. And the
<a href="https://github.com/egonw/chembl.rdf">(original) scripts to create RDF from the MySQL dump of ChEMBL</a> are also on GitHub.</p>

<p><img src="https://media.springernature.com/lw685/springer-static/image/art%3A10.1186%2F1758-2946-5-23/MediaObjects/13321_2012_Article_469_Figa_HTML.gif" alt="" /></p>

<p>This paper outlines the <a href="http://www.jcheminf.com/content/3/1/15">RDF</a> as it has evolved from various earlier projects. The above
diagram visualizes the basic structure (red), various Linked Data resources linked too (blue) and illustrates how various ontologies are used,
such as the <a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0025513">CHEMINF</a>, <a href="http://bibliontology.com/">BIBO</a>,
and <a href="http://www.jbiomedsem.com/content/1/S1/S6">CiTO</a> ontologies.</p>

<p>Additionally, various applications and links are described developed by various co-authors. For example, Peter worked on the use in
<a href="http://bio2rdf.org/">Bio2RDF</a> and Bin and David on <a href="http://cheminfov.informatics.indiana.edu:8080/">Chem2Bio2RDF</a>. Andra developed
an extension for his (#altmetric) <a href="http://citedin.org/">CitedIn</a> resource, giving credit to a paper when data in it is extracted into
ChEMBL. Ola, Valery, and Anthony developed a <a href="http://www.bioclipse.net/decision-support">Bioclipse Decision Support</a> extension,
which supports a nearest neighbor search in ChEMBL using <a href="http://chemspider.com/">ChemSpider</a>. Of course, Ola also hosts
<a href="http://rdf.farmbio.uu.se/chembl/snorql/">the SPARQL end point</a> of which you can monitor the uptime at the also cool
<a href="http://labs.mondeca.com/sparqlEndpointsStatus/details/farmbio-chembl.html">mondeca.com service</a>:</p>

<p><img src="/assets/images/mondecaUptime.png" alt="" /></p>

<p>(Yes, I think I have all the cool buzzwords covered in this paper. Sadly, marketing is needed nowadays as a scientist. Where is the
time that you could rant on page after page in all your domain specific jargon, not having to worry if your reader would understand
it immediately, or without a university degree…)</p>

<p>What this paper does not describe, is all the things I did with ChEMBL-RDF in the <a href="http://www.openphacts.org/">Open PHACTS</a> project
(<a href="https://twitter.com/open_phacts">@Open_PHACTS</a>), which includes the use of <a href="http://qudt.org/">QUDT</a> and the
<a href="https://github.com/egonw/jqudt">jQUDT</a> library for unit normalization outlined in <a href="http://www.bigcat.unimaas.nl/~egonw/units/">this document</a>
and the use of VoID for link sets as described in <a href="http://www.openphacts.org/specs/2012/WD-datadesc-20121019/">this document</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="chembl" /><category term="rdf" /><category term="cito" /><category term="cheminf" /><category term="doi:10.1186/1758-2946-5-23" /><category term="doi:10.1186/1758-2946-3-15" /><category term="ontology" /><category term="doi:10.1371/JOURNAL.PONE.0025513" /><category term="justdoi:10.1186/2041-1480-1-S1-S6" /><category term="chemspider" /><category term="openphacts" /><summary type="html"><![CDATA[]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/mondecaUptime.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/mondecaUptime.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The Molecular Chemometrics Principles #3: stand on shoulders</title><link href="https://chem-bla-ics.linkedchemistry.info/2010/08/14/molecular-chemometrics-principles-3.html" rel="alternate" type="text/html" title="The Molecular Chemometrics Principles #3: stand on shoulders" /><published>2010-08-14T00:00:00+00:00</published><updated>2010-08-14T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2010/08/14/molecular-chemometrics-principles-3</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2010/08/14/molecular-chemometrics-principles-3.html"><![CDATA[<p>I have blogged about two Molecular Chemometrics principles so far:</p>

<ul>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html">McPrinciple #1: access to data</a></li>
  <li><a href="https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html">McPrinciple #2: be clear in what you mean</a></li>
</ul>

<p>Peter’s post <a href="https://doi.org/10.59350/hphjc-qgr72">#solo10: Green Chain Reaction; where to store the data? DSR? IR? BioTorrent, OKF or ??? <i class="fa-solid fa-recycle fa-xs"></i></a>
gives me enough basis to write up a third principle:</p>

<p><strong>Molecular Chemometrics Principles #3</strong>: We make scientific progress if we build on past achievements.</p>

<p>Sounds logical, right? Practically, the way we share our cheminformatics knowledge makes this standing on shoulders pretty difficult.
But there is one particular aspect I would like to ask your attention for: you can contribute by making clear what shoulders
you would like to stand on. That is, where do you prefer to put your effort, and what message would you like to give to your user community.</p>

<p>In the aforelinked post, Peter asks where he should upload his data, and he suggest <a href="http://www.biotorrents.net/">BioTorrent</a> (see my review
<a href="https://chem-bla-ics.linkedchemistry.info/2010/04/18/bittorrents-for-science.html">BitTorrents for Science <i class="fa-solid fa-recycle fa-xs"></i></a>), DSpace, and <a href="http://www.ckan.net/">CKAN</a>.
Now, his <a href="http://www.google.se/search?sourceid=chrome&amp;client=ubuntu&amp;channel=cs&amp;ie=UTF-8&amp;q=%22Green+Chain+Reaction%22">Green Chain Reaction</a>
is picked up (see <a href="http://researchremix.wordpress.com/2010/08/11/green-chain-reaction-project-putting-my-minutes-where-my-mouth-is/">these</a>
<a href="http://scienceonlinelondon.wikidot.com/topics:green-chain-reaction">few</a> <a href="https://doi.org/10.59350/h2jq5-3np88">blog <i class="fa-solid fa-recycle fa-xs"></i></a> posts),
and the resulting data should be distributed as much as possible. The exact location does not really matter…</p>

<p>But…</p>

<p>By picking where you upload, you make a statement to your community: “<em>Look guys, we are distributing our data via Foo, because we believe those guys are doing good work! Perhaps you can support them too.</em>”.</p>

<p>This principle does not only apply to data, it applies to things too. For example, when
<a href="http://www.chemspider.com/blog/ichemlabs-and-rsc-chemspider-announce-partnership.html">iChemLabs and RSC ChemSpider Announce Partnership</a>
they do not just improve the user experience of ChemSpider (which I certainly won’t object against), but they also imply
“<em>Look dudes, your product is just not good enough and we do not want to help you improve it either</em>”.
Of course, ChemSpider has every right, and for them to succeed it is crucial to make decisions like this. Fortunately,
<a href="http://web.chemdoodle.com/installation.php">ChemDoodle is GPL</a>.</p>

<p>Every project with a user base has the opportunity to support shoulders, if they only visibly stand on them. By merely discussion the
<em>Green Chain Reaction</em>, I show to support this social web experiment. You can too. Use these powers wisely. May the McPrinciples be with you.</p>]]></content><author><name>Egon Willighagen</name></author><category term="mcprinciples" /><category term="solo10" /><category term="chemdoodle" /><category term="chemspider" /><category term="javascript" /><category term="justdoi:10.59350/hphjc-qgr72" /><category term="justdoi:10.59350/h2jq5-3np88" /><summary type="html"><![CDATA[I have blogged about two Molecular Chemometrics principles so far:]]></summary></entry><entry><title type="html">ChemSpider fail #1: SMILES</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/08/06/chemspider-fail-1-smiles.html" rel="alternate" type="text/html" title="ChemSpider fail #1: SMILES" /><published>2009-08-06T00:00:00+00:00</published><updated>2009-08-06T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/08/06/chemspider-fail-1-smiles</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/08/06/chemspider-fail-1-smiles.html"><![CDATA[<p>Cheminformatics is difficult, I know. But I thought I used a simple SMILES when I typed <em>C1CNCC1</em>, but <a href="http://www.chemspider.com/">ChemSpider</a>
got it wrong :) The correct structure should be <a href="http://en.wikipedia.org/wiki/Pyrrolidine">pyrrolidine</a>, not
<a href="http://en.wikipedia.org/wiki/Pyrrole">pyrrole</a>. I always mix up those names, so defaulted to ChemSpider to give me the correct name, which
<a href="http://www.chemspider.com/RecordView.aspx?rid=3fdc7226-0f87-44c5-b367-f3fdbda4bbde">ChemSpider knows</a> and where it also has the
SMILES correct… there just seems something wrong with there search dialog:</p>

<p><img src="/assets/images/chemSpiderFail.png" alt="" /></p>

<p>There has been <a href="http://www.chemspider.com/blog/?p=55">some talk about a ChemSpider Bugzilla</a>, but I don’t think this has
materialized yet, and I’ll have the default to <em>info-at-chemspider-dot-com</em> …</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemspider" /><category term="smiles" /><category term="inchikey:RWRDLPDLKQPQOW-UHFFFAOYSA-N" /><summary type="html"><![CDATA[Cheminformatics is difficult, I know. But I thought I used a simple SMILES when I typed C1CNCC1, but ChemSpider got it wrong :) The correct structure should be pyrrolidine, not pyrrole. I always mix up those names, so defaulted to ChemSpider to give me the correct name, which ChemSpider knows and where it also has the SMILES correct… there just seems something wrong with there search dialog:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/chemSpiderFail.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/chemSpiderFail.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">ChemSpider and the RSC: where next?</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/05/15/chemspider-and-rsc-where-next.html" rel="alternate" type="text/html" title="ChemSpider and the RSC: where next?" /><published>2009-05-15T00:00:00+00:00</published><updated>2009-05-15T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/05/15/chemspider-and-rsc-where-next</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/05/15/chemspider-and-rsc-where-next.html"><![CDATA[<p>Last Monday the <a href="http://www.indiana.edu/~cheminfo/network.html">CHMINF-L</a> brought the news to me that <a href="http://chemspider.com/">ChemSpider</a>
was acquired by the <a href="http://rsc.org/">RSC</a> (not the <a href="http://rsc.org/AboutUs/News/PressReleases/2009/ChemSpider.asp">press release</a>).
<a href="http://twitter.com/">Twitter</a> (<a href="http://twitter.com/egonwillighagen/statuses/1763364256">my Twitter post</a>) and
<a href="http://friendfeed.com/">FriendFeed</a> (see <a href="http://friendfeed.com/search?q=chemspider+rsc&amp;friends=egonw">this series</a>).</p>

<p>Reading blogs used to be to get the news, but this has changed. Still, blogging gives more freedom, more space. Blogs did soon
follow. <a href="http://www.steinbeck-molecular.de/steinblog/">Chris</a> was the first to
<a href="http://www.steinbeck-molecular.de/steinblog/index.php/2009/05/11/chemspider-bought-by-the-royal-society-of-chemistry/">blog about it</a>:</p>

<blockquote>
  <p>This is great news and I’m confident that it will be a move to even more openess in chemistry and cheminformatics.
It will also allow the RSC to use Tony fantastic tools for even more semantic markup of articles. I’m looking forward
to talking to everyone about the implications. For now, congratulations, Tony, and congratulations, RSC, for this
great deal.</p>
</blockquote>

<p>I think <a href="http://www.chemspider.com/blog/">Tony</a> himself <a href="http://www.chemspider.com/blog/the-royal-society-of-chemistry-acquires-chemspider.html">was next</a>:</p>

<blockquote>
  <p>This is good for us for a number of reasons. Specifically we will no longer have to deal with our very significant
resource limitations but more than that it lends credence and validation to the work that we have been doing over the
past 2 years. It seems so long ago now but ChemSpider was first unveiled to the world at the ACS Spring meeting 2007.
What began then only as a hobby project is now being recognized by the community as one of the primary resources for
internet chemistry.</p>
</blockquote>

<p>His network and insight in required data curation is what I think made ChemSpider a success.</p>

<p>Later views followed from <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=1891">Peter</a>, <a href="http://prospect.rsc.org/blogs/cw/?p=1829">Rich</a> and
<a href="http://blogs.nature.com/thescepticalchymist/2009/05/the_rsc_and_chemspider.html">Neil</a>. I have only congratulations,
which I hereby join, and expect that only future will tell us if our cheers are correct.</p>

<h2 id="where-next">Where next?</h2>
<p>As Tony indicated, the deal will practically mean better support for ChemSpider in terms of computing power, making if
easier for them to make upgrades, hence better uptime, etc. It may, indeed, also mean more data, provided from RSC archives,
as <a href="http://blogs.nature.com/thescepticalchymist/2009/05/the_rsc_and_chemspider.html">suggested by Neil</a>. More practically, I
can imagine seeing Project Prospect contributing <em>InChI-DOI</em> links to ChemSpider very soon.</p>

<p>And this would be one of the two recommendations I have to ChemSpider at this moment:</p>

<ol>
  <li>now linked to a publisher, and with both text mining efforts and expertise, focus on these InChI-DOI links, and, in
particular, focus on those InChI-DOI links which involve papers that describe measured properties of the molecules;</li>
  <li>with the increased support, finish the Open Data work done, by making it easy for people to download the
ChemSpider-OpenData subset. This, I believe, is crucial for a wider adoption in the OpenData community, as OpenData
which is practically made impossible to easily download is not Open enough. Previous priorities may have been focused
on setting up a viable commercial alternative, but with the RSC backing, this can no longer be a reason to not do this.</li>
</ol>

<p>Once more, congratulations to the ChemSpider-team and the involved RSC people, and very much looking forward to seeing
how this will change chemistry for the better!</p>]]></content><author><name>Egon Willighagen</name></author><category term="cheminf" /><category term="chemspider" /><category term="opendata" /><summary type="html"><![CDATA[Last Monday the CHMINF-L brought the news to me that ChemSpider was acquired by the RSC (not the press release). Twitter (my Twitter post) and FriendFeed (see this series).]]></summary></entry><entry><title type="html">Downloading Domoic Acid from PubChem</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/04/17/downloading-domoic-acid-from-pubchem.html" rel="alternate" type="text/html" title="Downloading Domoic Acid from PubChem" /><published>2009-04-17T00:10:00+00:00</published><updated>2009-04-17T00:10:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/04/17/downloading-domoic-acid-from-pubchem</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/04/17/downloading-domoic-acid-from-pubchem.html"><![CDATA[<p>The identity of <a href="http://en.wikipedia.org/wiki/Domoic_acid">domoic acid</a> has been under discussion (see
<a href="http://www.chemspider.com/blog/the-plot-thickens-on-domoic-acid.html">here</a>, <a href="http://www.chemspider.com/blog/where-does-ce-news-source-its-chemical-structures.html">here</a>
and <a href="http://www.chemspider.com/blog/providing-some-structured-support-with-chemspiders-wikipedia-services.html">here</a>).
(And I very much like the <a href="http://www.chemspider.com/">ChemSpider</a> service to make it easy to
<a href="http://www.chemspider.com/blog/providing-some-structured-support-with-chemspiders-wikipedia-services.html">copy data from ChemSpider into WikiPedia ChemBoxes</a>;
cheers!)</p>

<p>Now, my practical in next weeks <a href="https://apps.sourceforge.net/mediawiki/cdk/index.php?title=CDK_Workshop_2009">CDK Workshop will</a> use
<a href="http://groovy.codehaus.org/">Groovy</a> (please install it on your laptop!), and am hacking up example scripts for the course material,
and came up with this script to download the structure of <a href="http://pubchem.ncbi.nlm.nih.gov/summary/summary.cgi?cid=5282253">domoic acid</a>
from <a href="http://pubchem.ncbi.nlm.nih.gov/">PubChem</a> (CID:5282253):</p>

<script src="https://gist.github.com/97067.js"></script>]]></content><author><name>Egon Willighagen</name></author><category term="cdk" /><category term="pubchem" /><category term="chemspider" /><category term="wikipedia" /><category term="inchikey:VZFRNCSOCOPNDB-AOKDLOFSSA-N" /><summary type="html"><![CDATA[The identity of domoic acid has been under discussion (see here, here and here). (And I very much like the ChemSpider service to make it easy to copy data from ChemSpider into WikiPedia ChemBoxes; cheers!)]]></summary></entry><entry><title type="html">Open NMR data: raw curves and annotated peak lists</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/03/04/open-nmr-data-raw-curves-and-annotated.html" rel="alternate" type="text/html" title="Open NMR data: raw curves and annotated peak lists" /><published>2009-03-04T00:00:00+00:00</published><updated>2009-03-04T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/03/04/open-nmr-data-raw-curves-and-annotated</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/03/04/open-nmr-data-raw-curves-and-annotated.html"><![CDATA[<p>Games are known to trigger technical innovation. But recently it also triggered innovation on open chemical databases. Jean-Claude
<a href="http://usefulchem.blogspot.com/2009/03/spectral-game-update.html">reported</a>:</p>

<blockquote>
  <p>We are very excited by what we have put together so far. There are currently 457 H NMR, 389 C NMR, 11 IR and 29 NIR spectra. This
is only possible because of people who submitted their spectra to ChemSpider as Open Data - please keep uploading!</p>
</blockquote>

<p>Now, the <a href="http://nmrshiftdb.org/">NMRShiftDB</a> also hosts quite a number of NMR spectra, and I have a hobby to submit spectra,
particularly for rare nuclei. In particular, I think it is fun to to have as many as possible structures which have spectra for
all the nuclei in that structure. <a href="http://en.wikipedia.org/wiki/Benzene">Benzene</a> is a simple example for which NMR spectra are
available for all nuclei (see <a href="http://nmrshiftdb.chemie.uni-mainz.de/portal/js_pane/P-Results/nmrshiftdbaction/showDetailsFromHome/molNumber/7901">this entry</a>).</p>

<p>Now, the main difference between the NMRShiftDB and <a href="http://www.chemspider.com/">ChemSpider</a> spectral data is the the first are annotated
peak lists (each shift is assigned to an atom), and the latter are full, but unannotated, spectral curves. So, there are quite a few
things you could do here. For example, see which structures which NMR curves are not yet annotated in NMRShiftDB.
<a href="http://www.chemspider.com/blog/">Antony</a> pointed me to <a href="http://www.chemspider.com/spectra.aspx">this page</a> which is an overview
of all spectral data in ChemSpider, but that page is difficult to machine process. Partly, because it is a mix of Open and
Proprietary data, and partly because it uses JavaScript to navigate the table. (BTW, RDF interfaces to both resources would
be much more helpful, and simply allow me to query all molecules which have a spectrum which is Open, and which is not found
in the NMRShiftDB. I am working on a RDF interface to NMRShiftDB.)</p>

<p>Antony also <a href="http://usefulchem.blogspot.com/2009/03/spectral-game-update.html#c6864087472599346745">asked</a> me to advertise the
option to upload Open spectral curves to ChemSpider. So, hereby. However, I really do hope ChemSpider will make it easier for
others to reuse all the Open Data, as having to machine browsing the linked HTML interface is a waste of ChemSpider computing
resources.</p>

<p><strong>Update</strong>: the game is now available from <a href="http://spectralgame.com/">spectralgame.com</a>.</p>]]></content><author><name>Egon Willighagen</name></author><category term="opendata" /><category term="chemspider" /><category term="nmr" /><category term="nmrshiftdb" /><summary type="html"><![CDATA[Games are known to trigger technical innovation. But recently it also triggered innovation on open chemical databases. Jean-Claude reported:]]></summary></entry><entry><title type="html">YouTube for Chemistry</title><link href="https://chem-bla-ics.linkedchemistry.info/2009/02/04/youtube-for-chemistry.html" rel="alternate" type="text/html" title="YouTube for Chemistry" /><published>2009-02-04T00:00:00+00:00</published><updated>2009-02-04T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2009/02/04/youtube-for-chemistry</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2009/02/04/youtube-for-chemistry.html"><![CDATA[<p><a href="http://www.chemspider.com/">ChemSpider</a> has set up <a href="http://www.chemspider.com/blog/why-are-chemical-structures-like-youtube-videos.html">embeddable chemistry widget</a>
(per <a href="http://blog.openwetware.org/scienceintheopen/">Cameron</a>’s idea), much like <a href="http://youtube.com/">YouTube</a>. I just have to try that.
Unlike YouTube, you need to be registered and logged in to use the functionality (I hope the requirement will be dropped):</p>

<div class="language-html highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">&lt;script </span><span class="na">type=</span><span class="s">"text/javascript"</span> <span class="na">src=</span><span class="s">"http://www.chemspider.com/csjsapi.ashx?op=img&amp;amp;tk=3d178e75-a272-4d60-8ca9-5b1183a0e746&amp;amp;id=171&amp;amp;w=120&amp;amp;p=1&amp;amp;eid=%22azijnzuur%22"</span><span class="nt">&gt;&lt;/script&gt;</span>
</code></pre></div></div>

<p>There is an option to have ChemSpider link back to blog, and I will have to figure out how to enable
<a href="http://cb.openmolecules.net/">Chemical blogspace</a> to extract the InChI from the underlying JavaScripts.</p>

<p><strong>Update</strong>: I noticed that the ChemSpider server was a bit sluggish this morning, and that loading my blog page halts at loading the
JavaScript… Tony, I suggest to use some Ajax magic here, with a really fast JavaScript download (using an almost static bit of
JavaScript), and then a Ajax to access to slower bits, which might involve image generation and database lookup.</p>

<p><strong>Update2</strong>: the feature was already under development before Cameron asked about it.</p>

<p><strong>Update3</strong>: the script is no longer working, and I made the code visible instead, for historic reasons.</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemspider" /><summary type="html"><![CDATA[ChemSpider has set up embeddable chemistry widget (per Cameron’s idea), much like YouTube. I just have to try that. Unlike YouTube, you need to be registered and logged in to use the functionality (I hope the requirement will be dropped):]]></summary></entry><entry><title type="html">Open{Data|Source|Standards} is not enough: we need Open Projects</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough.html" rel="alternate" type="text/html" title="Open{Data|Source|Standards} is not enough: we need Open Projects" /><published>2008-11-07T00:00:00+00:00</published><updated>2008-11-07T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/11/07/opendatasourcestandards-is-not-enough.html"><![CDATA[<p>The <a href="http://blueobelisk.sourceforge.net/wiki/Main_Page">Blue Obelisk</a> mantra <a href="http://blueobelisk.sourceforge.net/wiki/ODOSOS">ODOSOS</a>,
Open Data, Open Source, Open Standards, is well known, and much cited too. <a href="http://usefulchem.blogspot.com/">Jean-Claude Bradley</a>
popularized the <a href="http://en.wikipedia.org/wiki/Open_Notebook_Science">Open Notebook Science</a> (ONS). This has always been nagging me a bit,
because the <a href="http://cdk.sf.net/">CDK</a>, <a href="http://www.jmol.org/">Jmol</a>, JChemPaint and other chemistry projects have done that for much
longer, though we did not use notebooks as much, so called it just an open source project. It really is no different, IMO, though
surely, there are differences.</p>

<p>Anyway, the key thing which ONS and CDK and Jmol share, is that they use an Open Notebook. Not every Open Source or Open Data project does.
Actually, many scientific Open Source are not open Projects! They are more like the Cathedral than the wished-for Bazaar (see
<a href="http://en.wikipedia.org/wiki/The_Cathedral_and_the_Bazaar">The Cathedral and the Bazaar</a>). So, Open Source (science) projects are certainly not ONS projects by default!</p>

<p>Now, the CDK actually is ONS, it is a Bazaar. The notebooks we use include:</p>

<ul>
  <li>open project via <a href="https://sourceforge.net/mail/?group_id=20024">mailing lists</a></li>
  <li>open methods/results via <a href="https://sourceforge.net/svn/?group_id=20024">subversion</a></li>
  <li>informal reporting via blogs (e.g. <a href="http://rguha.wordpress.com/">Rajarshi</a>, <a href="http://www.steinbeck-molecular.de/steinblog/">Christoph</a>, <a href="http://cdktaverna.wordpress.com/">Thomas</a>, mine)</li>
  <li>informal reporting via <a href="http://www.cdknews.org/">CDK News</a></li>
</ul>

<p>What more would you wish for? That’s not a rhetorical question. Remember that every reader of this blog is in
<a href="https://chem-bla-ics.linkedchemistry.info/2007/11/27/be-in-my-advisory-board-1-being-good.html">my advisory board <i class="fa-solid fa-recycle fa-xs"></i></a>!</p>

<p>Unfortunately, I do not create work at a workbench myself, so I do not produce new knowledge myself, other than extracted from existing
data. That’s really a shame, and I really do hope that Jean-Claude or <a href="http://blog.openwetware.org/scienceintheopen">Cameron</a> will send
me a box to measure solubilities (see <a href="http://usefulchem.blogspot.com/2008/10/rdf-triples-for-open-notebook-science.html">here</a>,
<a href="http://usefulchem.blogspot.com/2008/11/ons-solubility-web-query.html">here</a>, and
<a href="http://anybody.cephb.fr/perso/lindenb/tmp/jcbradley.rdf">here</a>,
<a href="http://rguha.wordpress.com/2008/11/06/solubility-queries-and-the-google-visualization-api/">here</a> for first data exploration),
even though I cannot participate in the <a href="http://usefulchem.blogspot.com/2008/11/submeta-open-notebook-science-awards.html">challenge</a>.
(hint, hint :)</p>

<h2 id="from-cathedral-to-bazaar-in-life-sciences">From Cathedral to Bazaar in Life Sciences</h2>

<p>One Cathedral we ran into with <a href="http://www.bioclipse.net/">Bioclipse</a> was <a href="http://www.biocatalogue.org/">BioCatalogue</a>,
which will serve as website where people can annotate and categorize (web) services. While the project has been around for a while, the
website was rather uninformative. Fortunately, the projects is going to open up, and be more Bazaar-like. For example, they
now started a <a href="http://www.biocatalogue.org/wiki">wiki</a> and a
<a href="http://listserv.manchester.ac.uk/cgi-bin/wa?SUBED1=biocatalogue-friends&amp;A=1">mailing list</a>. I hope these efforts will continue,
so that I can contribute from my point of view!</p>

<p>The <a href="http://embraceregistry.net/">EMBRACE Registry</a> is a project with similar goals and a rather nice outcome (which I learned about on
<a href="https://chem-bla-ics.linkedchemistry.info/2008/11/03/embrace-workshop-in-uppsala.html">Monday <i class="fa-solid fa-recycle fa-xs"></i></a>). It is actually anticipate to be replaced by or merge
with BioCatalogue. So, all data I entered, <a href="http://prints.cs.man.ac.uk:8081/category/tags/cheminformatics">cheminformatics workflows</a>
(look, <a href="https://chem-bla-ics.linkedchemistry.info/2008/10/18/chemoinformatics-p0wned-by.html">no ‘o’ <i class="fa-solid fa-recycle fa-xs"></i></a>), will later be available from BioCatalogue too.
That is already my first contribution to BioCatalogue. One enormously interesting feature of the Registry, is that is allows uploading of
code to test the service. This will mean the Registry will not only poll if the service is still online (by checking the WSDL file), it
will also test if the service behaves properly. Now, immediate thoughts are mashups with <a href="http://www.myexperiment.org/">MyExperiment</a>.
Each WSDL entry in the Registry points to MyExperiment workflows that use them, and the workflow page would indicate the status of all
used WDSL services. This integration was already anticipated long before I thought about it, as the involved Cathedrals were nicely
located in the same floor in Manchester.</p>

<p>Below is a screenshot from the EMBRACE Registry for the <a href="http://www.chemspider.com/">ChemSpider</a>
<a href="http://prints.cs.man.ac.uk:8081/service/massspecapi">WDSL entry</a> for <a href="http://www.myexperiment.org/workflows/97">a workspace</a>
I <a href="https://chem-bla-ics.linkedchemistry.info/2007/11/26/metabolomics-workflows-in-taverna.html">uploaded <i class="fa-solid fa-recycle fa-xs"></i></a> about a year ago to MyExperiment:</p>

<p><img src="/assets/images/registry.png" alt="" /></p>

<p>BTW, ChemSpider has an Advisory Board of which I am member, but it is also a classical (and intentional) Cathedral project. We do share common interests though, which makes us collaborate.</p>

<h2 id="why-important">Why Important?</h2>

<p>One recurrent theme in Open Source is <a href="http://en.wikipedia.org/wiki/Given_enough_eyeballs">given enough eyeballs, all bugs are shallow</a>.
This surely applies to science as well. The difference between the two is that in current science the eyes only inspect with a delay of at
least 6 months. Current practice is that research is finished (delay), and when decided publishable written up a paper (delay, and loosing
valuable information in the process, as you can read in my blog all the time), and published (even more delay). ONS changes that, and so do
Bazaar-like open source projects, such as the CDK, Jmol and Bioclipse. They bugs are present, whether we like it or not, not just in source
code, but in science too. Theories get overthrown, but why should we like the long delays current scientific good practice? Hate it! Work
around it. Use the Bazaar. Use ONS!</p>

<p>Now, ONS actually needs Open Source, allowing them to deal effectively with the data they produce; to allow extraction of new scientific
knowledge from the measurements. If Rajarshi and Pierre would not have made their efforts, other could not easily join in, leading to
those much hated delays. Bugs should be shallow, and openness allows us to make those bugs visible. We can prove that there is a bug,
without having to reproduce data ourselves, leading to those nasty delays again. Just copy the data, compare it to your own, do your
analysis.</p>

<p>One recent project in open source chemistry dealing with making bugs visible, is the web page set up by Andreas Tille for the
<a href="http://alioth.debian.org/projects/debichem">DebiChem project</a>. His page <a href="http://cdd.alioth.debian.org/debichem/bugs/">summarizes the bugs</a>
listed for the chemistry in Debian (which includes the Blue Obelisk projects <a href="http://packages.debian.org/lenny/avogadro">Avogadro</a>,
<a href="http://packages.debian.org/lenny/bodr">BODR</a>, <a href="http://packages.debian.org/lenny/libcdk-java">CDK</a>,
<a href="http://packages.debian.org/lenny/chemical-mime-data">Chemical MIME Data</a>,
<a href="http://packages.debian.org/lenny/kalzium">Kalzium</a> and <a href="http://packages.debian.org/lenny/openbabel">OpenBabel</a>):</p>

<p><img src="/assets/images/debichem.png" alt="" /></p>

<p>This data analysis helps the projects being analyzed.</p>

<h2 id="packaging">Packaging</h2>

<p>This brings me to a last topic, for this blog: packaging using Open Standards. In order to allow those eyeballs to spot bugs, it is of the
utmost importance to package your results in Open Standards, and not just one, but likely many. For Open Source projects this ultimately
means Distribution Packages (deb or rpm). If that goal has been achieved, you know your results can be read by anyone. Software should be
installable (make, ant, cmake, etc), and Data should be readable (no PDF, but RDF, XML, JSON, or whatever standard). Preferably not Excel,
as this is too free format (as Rajarshi also <a href="http://rguha.wordpress.com/2008/11/06/solubility-queries-and-the-google-visualization-api/">indicated</a>),
but with some added conventions it may do well. Blue Obelisk project are generally doing well in terms of packaging.</p>

<p>For the CDK, which already is reasonably well packaged, I am currently working on <a href="http://cdk.svn.sourceforge.net/viewvc/cdk/cdk-eclipse/trunk/">Eclipse</a>
and <a href="http://cdk.svn.sourceforge.net/viewvc/cdk/cdk-pom/trunk/">Maven2</a> packages. The former is already being used by Bioclipse, while the
second aims at <a href="https://sourceforge.net/projects/cml">Jumbo</a> (which has just seen a
<a href="https://sourceforge.net/project/showfiles.php?group_id=51361">new release</a>. <a href="http://wwmm.ch.cam.ac.uk/blogs/downing/">Jim</a>,
I’m happy to see the CMLDOM/Jumbo split!), <a href="http://www.cdk-taverna.de/">CDK-Taverna</a>, and possibly a third (Paula, what for do you plan
to use it?). The POM export is not fully working yet, but with four research sites involved in this Open Project, I’m sure we’ll work
it out.</p>

<p>The bottom line is, scientific progress would benefit so much from a Bazaar approach. And the key thing is not collaboration; that’s
something you can do in a Cathedral-like fashion too. No, the key thing is to be Open and allow anyone, even your worst nightmare, to
comment on what you do. Let him prove you wrong, openly, that is.</p>

<p>OK, there it is. My open notebook entry for this week. Now you know what I have been up to this week.</p>]]></content><author><name>Egon Willighagen</name></author><category term="odosos" /><category term="chemspider" /><category term="workflow" /><category term="cdk" /><category term="bioclipse" /><category term="cml" /><category term="debian" /><category term="eclipse" /><category term="rdf" /><category term="jmol" /><category term="blue-obelisk" /><summary type="html"><![CDATA[The Blue Obelisk mantra ODOSOS, Open Data, Open Source, Open Standards, is well known, and much cited too. Jean-Claude Bradley popularized the Open Notebook Science (ONS). This has always been nagging me a bit, because the CDK, Jmol, JChemPaint and other chemistry projects have done that for much longer, though we did not use notebooks as much, so called it just an open source project. It really is no different, IMO, though surely, there are differences.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/registry.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/registry.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Does ChemSpider really violate Open Data with CC SA?</title><link href="https://chem-bla-ics.linkedchemistry.info/2008/05/10/does-chemspider-really-violate-open.html" rel="alternate" type="text/html" title="Does ChemSpider really violate Open Data with CC SA?" /><published>2008-05-10T00:00:00+00:00</published><updated>2008-05-10T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2008/05/10/does-chemspider-really-violate-open</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2008/05/10/does-chemspider-really-violate-open.html"><![CDATA[<p><a href="http://www.chemspider.com/">ChemSpider</a> <a href="http://www.chemspider.com/blog/it-appears-chemspider-does-bad-by-using-creative-commons-licenses.html">is afraid</a>
they are doing something bad because they release their data as <a href="http://creativecommons.org/licenses/by-sa/3.0/">CC-BY-SA</a>.
Because, John Wilbanks says in Peter’s blog:</p>

<blockquote>
  <p>I would add to it that I’d like to see a meaningful discussion of the
risks of Share Alike and Attribution on <strong>data integration</strong>. Chemspider’s
move to CC-BY-SA fits into this discussion nicely - it’s a total
violation of the open data protocol we laid out at SC, which says “Don’t
Use CC Licenses on Data” - <strong>but it does conform inside the broader OKD.</strong></p>
</blockquote>

<p>Now, let’s take this into pieces.</p>

<ol>
  <li>John notes that ChemSpider is in compliance with the <a href="http://www.opendefinition.org/1.0/">OKD</a>. This means, that ChemSpider thinks
about Open Data just like the <a href="http://en.wikipedia.org/wiki/Open_Knowledge_Foundation">Open Knowledge Foundation</a> does. I’ve scanned
through the OKD, and it indeed seems to support the BY and SA clauses of the CC. So, Chemspider did not do a bad thing.</li>
  <li>Data integration is tricky: you have to keep track of license information on an entry-by-entry level. For each fact, you keep to track the
source, and associate the source with it’s original license. For example, the <a href="http://www.nmrshiftdb.org/">NMRShiftDB</a>
information in ChemSpider should be <a href="http://www.gnu.org/copyleft/fdl.html">GNU FDL</a>.</li>
  <li>OpenX licenses may be viral. This holds for the <a href="http://www.gnu.org/licenses/gpl.html">GNU GPL</a> as well as for the CC-BY-SA.
Nothing new there. It just requires that when you would like to incorporate the ChemSpider data into a larger database, that database
has to be CC-BY-SA too, or likely at least CC-SA.</li>
</ol>

<p>Summarizing, I think ChemSpider did a good thing, and that ChemSpider does <strong>not</strong> violate the OpenData idea, but instead, that the CC-BY-SA and
the OKD violates John’s requirements for integrating data resources (apparently based on a two year legal study). That has nothing to do with ChemSpider.</p>

<p>Now, people will always have different opinions on Openness. The original BSD clause had a
<a href="http://en.wikipedia.org/wiki/BSD_License#UC_Berkeley_advertising_clause">restrictive ‘advertisement’ clause</a>, not Open enough for at least the
<a href="http://www.debian.org/social_contract#guidelines">Debian Free Software Guidelines</a> (DFSG), while still open source. The clause was
later removed from the BSD license.</p>

<p>Another <a href="http://www.debian.org/">Debian</a> example is Firebox, which is named <a href="http://packages.debian.org/iceweasel">IceWeasel</a> in Debian,
because the ‘license’ on the Firefox name is not open enough.</p>

<p>Another problem with the definition of Openness, is the viral aspect of some licenses (see earlier). For some, the GPL is not open enough,
because it does not give people the freedom to license their software they like themselves, something the BSD and MIT licenses do allow.
There is ongoing debate (and that should be ongoing) on how much <em>freedom</em> a license must provide to be called Open. The whole OpenAccess
discussion is similar (see e.g. <a href="http://www.google.com/search?q=strong+weak+open+access+site%3Awwmm.ch.cam.ac.uk&amp;btnG=Search">Peter’s story on this</a>),
where the discussion on the minimal amount of freedom is even worse.</p>

<p>Should we worry about ChemSpider being ‘only’ CC-BY-SA? Maybe. Data is not software, but I disagree that viral license would be OK for software, but NOT for data. That’s just BSD-versus-GPL all over again. I am happy about OpenBabel being GPL, and I am happy about ChemSpider being CC-BY-SA too.</p>

<p>All that said, these discussion are important. And creating good definitions of what freedoms are required, are crucial in deciding whether something is Open. The Blue Obelisk does not have/use such definitions yet, and we should start discussing this, and define a Blue Obelisk ODOSOS Guidelines. Please no funny jokes about how we can boogy then :)</p>

<p>Now, looking forward to hearing what you think about these issues… Looking forward to the other blog items!</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemspider" /><category term="copyright" /><category term="nmrshiftdb" /><summary type="html"><![CDATA[ChemSpider is afraid they are doing something bad because they release their data as CC-BY-SA. Because, John Wilbanks says in Peter’s blog:]]></summary></entry></feed>