<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemistry.xml" rel="self" type="application/atom+xml" /><link href="https://chem-bla-ics.linkedchemistry.info/" rel="alternate" type="text/html" /><updated>2026-07-18T13:36:15+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/feed/by_tag/chemistry.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">The TDCC NES Col-Lab Retreat</title><link href="https://chem-bla-ics.linkedchemistry.info/2026/02/21/the-tdcc-nes-col-lab-retreat.html" rel="alternate" type="text/html" title="The TDCC NES Col-Lab Retreat" /><published>2026-02-21T00:00:00+00:00</published><updated>2026-02-21T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2026/02/21/the-tdcc-nes-col-lab-retreat</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/02/21/the-tdcc-nes-col-lab-retreat.html"><![CDATA[<p>Last autumn two TDCC projects started, <em>FAIR4ChemNL</em> (<a href="https://chem-bla-ics.linkedchemistry.info/2026/02/08/open-infrastructures.html">with the PeerTube channel</a>
and doi:<a href="https://doi.org/10.61686/XVYQV45374">10.61686/XVYQV45374</a>) and <em>FAIRify for metabolomics data</em>
(doi:<a href="https://doi.org/10.61686/CSGIP04334">10.61686/CSGIP04334</a>). But I haven’t written much on either yet and what the role is our research group in these projects.</p>

<p>Let’s start with what the TDCC actually are: they are <a href="https://tdcc.nl/">Thematic Digital Competence Centres</a>:</p>

<blockquote>
  <p>The Thematic Digital Competence Centres (TDCCs) are network-based initiatives set up by NWO and the Dutch academic
community to broker investments into research data management projects. The three TDCCs are national and discipline
based, with one pillar each for the Social Sciences &amp; Humanities (SSH), Natural and Engineering Sciences (NES) and
Life Sciences &amp; Health (LSH). The networks will help formulate and facilitate projects designed to promote the adoption
of open data, software and research practices, alongside the development of the necessary expertise.</p>
</blockquote>

<p>So, where initiatives like <a href="https://www.go-fair.org/">GO FAIR</a> had centers of competencies (the implementation networks),
they did not have funding for them. This was a main reason why the <em>Chemistry Implementation Network</em> (ChIN,
doi:<a href="https://doi.org/10.1162/dint_a_00035">10.1162/dint_a_00035</a>) did not take off.
The TDCCs do not provide a lot of money, but enough to support disseminating expertise and promote some key ideas.</p>

<p>The idea is that combined with other efforts, it strengthens the level of FAIR in the Dutch research community.
I have to say, this is much needed, as the level of FAIR data in journal publications is so much to wish for,
and still mostly absent.</p>

<p>The FAIR4ChemNL project already had a networking activity during the writing of the proposal, the workshop already
back in 2024 that I <a href="https://chem-bla-ics.linkedchemistry.info/2024/06/10/two-meetings.html">blogged about earlier</a>
(see also <a href="https://doi.org/10.5281/zenodo.15050550">this report</a>).
The FAIRify project is coordinated by the group that was key in the <em>Netherlands Metabolomics Center</em> (NMC), now the
<a href="https://metabolomicscentre.nl/">BeneLux Metabolomics Center</a>. During a postdoc at the NMC during my Wageningen
days, we already did a lot of FAIR competency building with <a href="https://chem-bla-ics.linkedchemistry.info/tag/metware">the MetWare project</a>.</p>

<h2 id="the-col-lab-retreat">The Col-Lab Retreat</h2>

<p>The <a href="https://tdcc.nl/about-tddc/nes/">TDCC-NES</a> organized a networking event in August last year,
the 2025 <a href="https://nescollab.nl/">TDCC-NES Col-Lab Retreat</a>. I am late with
reporting on it, but there simply was too much project management that took priority. The meeting was in the
wonderful Dutch town Schoorl, and the location is great for collaborative meetings. I had been there a year
earlier for an Open Science Retreat and was happy to go back.</p>

<p>During the unconference-style meeting <a href="https://tdcc.nl/creating-space-for-our-community-the-story-of-our-nes-col-lab-retreat/">various topics were discussed</a>
in breakout groups, and because of the two TDCC projects, I was particularly interested in the <em>Metadata and interoperability</em>
topic. Partly because this is how we can make eletronic lab notebooks automatically push metadata to
registries (and <a href="https://www.linkedin.com/in/rory-macneil-68a80011/">Rory Macneil</a> was also in Schoorl,
of <a href="https://www.researchspace.com/">RSpace/ResearchSpace</a> which already integrated with various open
platforms), and partly because I wanted to continue explore <a href="https://chem-bla-ics.linkedchemistry.info/tag/nanopub">nanopublications</a>
with <a href="https://fediscience.org/@rupdecat">Christian Meesters</a>, which could be the envelope to distribute
the metadata. For the last, I was looking at the Java library for nanopublications
(see <a href="https://github.com/Nanopublication/nanopub-java/pull/52">this PR</a>.</p>

<p>The idea that ELNs automatically share metadata about experiments is something that is attractive.
It would require no involvement from the researcher, would be fully automatic, and drive interest
(users, peer reviewers) to experiments and experimental data. Something that is still absurdly hard
is to do a search for experiments that measured the melting point of some chemical. How
awesome would it be if ELNs would automatically register chemicals from the experiment in,
for example, <a href="https://pubchem.ncbi.nlm.nih.gov/">PubChem</a>.</p>

<p>We had the idea of applying for a Lorentz Workshop, but the earliest deadline was too early, but
maybe it is time to pick up that idea again. Interoperability standards already exist, like
the aforementioned nanopubs, but also <a href="https://www.researchobject.org/ro-crate/">RO-Crates</a> that are also studied by Jente Houweling
in the VHP4Safety project (see <a href="https://platform.vhp4safety.nl/data">this Data tab</a> for a preview).</p>]]></content><author><name>Egon Willighagen</name></author><category term="fair" /><category term="doi:10.1162/DINT_A_00035" /><category term="chemistry" /><category term="metabolomics" /><category term="fair4chemnl" /><category term="fairify" /><category term="cito:citesAsEvidence:10.5281/ZENODO.15050550" /><category term="nanopub" /><category term="crate" /><category term="pubchem" /><summary type="html"><![CDATA[Last autumn two TDCC projects started, FAIR4ChemNL (with the PeerTube channel and doi:10.61686/XVYQV45374) and FAIRify for metabolomics data (doi:10.61686/CSGIP04334). But I haven’t written much on either yet and what the role is our research group in these projects.]]></summary></entry><entry><title type="html">25 years of the Chemistry Development Kit</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/09/28/25-years-of-the-chemistry-development-kit.html" rel="alternate" type="text/html" title="25 years of the Chemistry Development Kit" /><published>2025-09-28T00:00:00+00:00</published><updated>2025-09-28T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/09/28/25-years-of-the-chemistry-development-kit</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/09/28/25-years-of-the-chemistry-development-kit.html"><![CDATA[<p>Twenty five years ago the <a href="https://cdk.github.io/">Chemistry Development Kit</a> (CDK) was founded. The Chemistry and Internet (<a href="https://www.google.com/search?q=ChemInt2000">ChemInt2000</a>)
had just ended (it ran from 23 to 26 September) and my friend and I had taken the Amtrak night train from Washington to South Bend. At that time there
were two leading Java applets for chemistry, <a href="https://jchempaint.github.io/">JChemPaint</a> and <a href="http://jmol.org/">Jmol</a>. I had hacked Chemical Markup
Language support into both of them, and <a href="https://chemistry.nd.edu/people/dan-gezelter/">Dan Gezelter</a> (Jmol and <a href="https://openscience.org/">openscience.org</a>),
<a href="http://www.steinbeck-molecular.de/steinblog/">Christoph Steinbeck</a> (JChemPaint), and me took the opportunity of being in North America
to discuss if we could use a common code base. Chris’ <em>compchem</em> had done something similar. Peter Murray-Rust, who had also attended ChemInt2000
like me and Chris did not attend.</p>

<p>I do not remember exactly, but I guess we must have met on the 28th and 29th? Maybe already on Wednesday. During this meeting we discussed a common
data model (yes, Jmol used the CDK data model at some point) and somewhere during the meeting we wrote down a name for the project. There was the
Java Development Kit, so this could be the Chemistry Development Kit. The name stuck.</p>

<p>A quick post like this cannot do credit to the history of the CDK, nor of everyone involved in the past or still is. You can browse some of the history
of the CDK in <a href="https://chem-bla-ics.linkedchemistry.info/tag/cdk">my blog</a> and in <a href="http://www.steinbeck-molecular.de/steinblog/index.php/category/chemistry-development-kit/">Chris’ blog</a>.
It has been an amazing journey and with a small grant just behind us (with  Alyanne de Haan, René van der Ploeg, and Marc Teunis from Hogeschool Utrecht),
and all the awesome things ongoing (new JChemPaint, various extensions, upgraded downstream tools), the CDK is alive and kicking.</p>

<p>A huge congrats and thanks to everyone (and every company and organization) who contributed code to the CDK with this huge milestone. There are a few people
that I want to particularly thank (see the AUTHORS file for all names): Chris, who in the late nineties made a difference with open source in chemistry,
Dan, for Jmol and hosting this memorable meeting at Notre Dame University, Rajarshi Guha, who operated <em>CDK Nightly</em> for many years, well before Travis
and Google Actions, Stefan, Miguel, Gilleain, and Christian, for many years of contributions to the CDK, and John Mayfield, the current
CDK release manager.</p>]]></content><author><name>Egon Willighagen</name></author><category term="cdk" /><category term="jchempaint" /><category term="jmol" /><category term="openscience" /><category term="chemistry" /><summary type="html"><![CDATA[Twenty five years ago the Chemistry Development Kit (CDK) was founded. The Chemistry and Internet (ChemInt2000) had just ended (it ran from 23 to 26 September) and my friend and I had taken the Amtrak night train from Washington to South Bend. At that time there were two leading Java applets for chemistry, JChemPaint and Jmol. I had hacked Chemical Markup Language support into both of them, and Dan Gezelter (Jmol and openscience.org), Christoph Steinbeck (JChemPaint), and me took the opportunity of being in North America to discuss if we could use a common code base. Chris’ compchem had done something similar. Peter Murray-Rust, who had also attended ChemInt2000 like me and Chris did not attend.]]></summary></entry><entry><title type="html">PFAS in the blood of the Dutch population</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population.html" rel="alternate" type="text/html" title="PFAS in the blood of the Dutch population" /><published>2025-07-06T00:00:00+00:00</published><updated>2025-07-06T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/07/06/pfas-in-the-blood-of-the-dutch-population.html"><![CDATA[<p>A recent report by the Dutch <a href="https://www.rivm.nl/">RIVM</a>, <em>PFAS in the blood of the Dutch population</em>
(doi:<a href="https://www.rivm.nl/bibliotheek/rapporten/2025-0094.pdf">10.21945/RIVM-2025-0094</a>), writes
that seven <a href="https://scholia.toolforge.org/chemical-class/Q648037">PFAS</a> compounds are found in blood samples
of all tested people. Another nine compounds are found in at least 1-in-10 people.
Because there is relevant data in the report on the 28 studied PFAS compound, I wanted to
have the report more FAIR than it is on the website. Why this report? Well, the chemistry and the
history is fascinating and brutal (I like <a href="https://www.youtube.com/watch?v=SC2eSujzrUY">this Veritasium video</a>).</p>

<p>The history tells me that our society may sound woke and leftish, in reality it is a continous fight
for basic human rights. (Something that plenty have been saying for years.)
In this case, a healtht life is the human right.</p>

<p>So, what can I do to make this report more FAIR?</p>

<h2 id="findable-in-wikidata">Findable in Wikidata</h2>

<p>Since this report has been <a href="https://news.google.com/search?q=PFAS%20in%20the%20blood%20of%20the%20Dutch%20population&amp;hl=en-US&amp;gl=US&amp;ceid=US%3Aen">mentioned in the news</a>,
it clearly is notable. The simplest thing to do is thus to just add it <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a>.
Because the DOI of the report had not been recorded yet, I could not let <a href="https://scholia.toolforge.org/">Scholia</a>
do it for me. But doing it manually is only a bit more work: <a href="https://www.wikidata.org/wiki/Q135222054">Q135222054</a>.
The provided metadata <a href="https://www.rivm.nl/en/news/first-nationwide-study-into-pfas-in-blood">on the RIVM website</a>
is minimal.</p>

<p>But we can do more. Particularly, because I want people to find this report when they look info knowledge
about the 28 studied chemicals, I added <a href="https://www.wikidata.org/wiki/Q135222054#P921">main subject</a> annotation
using the information in <em>Table 1</em> in the report. Using Scholia and the CAS registry number in the table,
I crosscheck the information in Wikidata is consistent with the report (and visa versa). It was.
I then added the Dutch name and acronym for most of them. Some already had the name as in the Table.
That gives us a nice “Topic scores” plot for <a href="https://scholia.toolforge.org/work/Q135222054">the Scholia page of the report</a>:</p>

<p><img src="/assets/images/pfas_report.png" alt="" /></p>

<p>The central PFAS bubble is also only one <em>main subject</em> but larger because many the specific PFAS compounds
are subclassing PFAS. And you may also note many smaller bubbles. These actually come from <em>main subject</em>
annotations of articles cited from the report. Because I added a few of them too. Not all, because many are
not in Wikidata (yet).</p>

<h2 id="findable-in-wikipathways">Findable in WikiPathways</h2>

<p>But since 16 of these compounds are readily found in human blood samples, that is handy knowledge when
doing metabolomics (on blood samples). Or (and I leave that to later blog post), we can map the experimental
data for Dordrecht versus the rest of The Netherlands to the PFAS compounds. That is relevant to research
by <a href="https://vhp4safety.nl">VHP4Safety</a>. There are many ways to see if you have PFAS in your dataset,
but since we have many controlled lists of genes in metabolites, I added one for common PFAS in human
blood samples. Well, the 16 common in Dutch blood samples:</p>

<p><img src="/assets/images/pfas_wikipathways.png" alt="" /></p>

<p>Each <em>metabolite</em> here is annotated with their Wikidata identifier, allowing us to map experimental
data on top of it. And we get links out to other databases almost for free:</p>

<p><img src="/assets/images/pfas_wikipathways_outlinks.png" alt="" /></p>

<p>And the link to Wikidata actually links to Scholia, so for the PFOA in the above example,
we can quickly see the boiling point, decomposition point, and melting point of this PFAS.
And literature with undoubtedly even more knowledge about this PFAS:</p>

<p><img src="/assets/images/pfas_scholia.png" alt="" /></p>

<p>Now, these two steps were mostly manual: drawing <a href="https://classic.wikipathways.org/index.php/Pathway:WP5579">WP5579</a>
in WikiPathways and adding the report annotations (<em>main subject</em> and <em>cites</em>) in Wikidata.</p>

<h2 id="findable-in-the-vhp4safety-compound-wiki">Findable in the VHP4Safety Compound Wiki</h2>

<p>As part of the VHP4Safety project, I am collecting information on chemicals studied in the context
of toxicology, safety, and risk assessment. Often specific collections of compounds studied as a whole.
This report is such a collection and provides experimental data on these compounds. So, I want this
report to be findable for the <a href="https://compoundcloud.wikibase.cloud/">VHP4Safety Compound Wiki</a> too.
Creating the collection is a manual step: <a href="https://compoundcloud.wikibase.cloud/wiki/Item:Q5145">Q5145</a>.</p>

<p>Now, because both Wikidata and our VHP4Safety Compound Wiki (a Wikibase instance) are semantic and support, I can use SPARQL
to create instructions to link the 28 compounds to the new collection. Now, arguably, that can be
done manually too, and maybe faster, for larger collections this is harder. So, I dug up
<a href="https://compoundcloud.wikibase.cloud/wiki/User:Egonw">my earlier notes</a> and got some useful
things together.</p>

<p>This query lists all 28 PFAS linked to the report <a href="https://w.wiki/Eepm">in Wikidata</a>:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="nv">?pfasLabel</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q135222054</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P921</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?pfas</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P31</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q113145171</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">SERVICE</span><span class="w"> </span><span class="nn">wikibase</span><span class="o">:</span><span class="ss">label</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nn">bd</span><span class="o">:</span><span class="ss">serviceParam</span><span class="w"> </span><span class="nn">wikibase</span><span class="o">:</span><span class="ss">language</span><span class="w"> </span><span class="s2">"[AUTO_LANGUAGE],mul,en"</span><span class="p">.</span><span class="w"> </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>Using federation powers, I can use this for <a href="https://edu.nl/ar9wf to match these up with our Wikibase">a SPARQL query</a>,
and return the results in QuickStatements that say <em>this VHP compound is part of the VHP collection</em>:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wb</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;https://compoundcloud.wikibase.cloud/entity/&gt;</span><span class="w">
</span><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wbt</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;https://compoundcloud.wikibase.cloud/prop/direct/&gt;</span><span class="w">

</span><span class="k">SELECT</span><span class="w"> </span><span class="p">(</span><span class="nb">SUBSTR</span><span class="p">(</span><span class="nb">STR</span><span class="p">(</span><span class="nv">?cmp</span><span class="p">),</span><span class="mi">45</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?qid</span><span class="p">)</span><span class="w"> </span><span class="nv">?P21</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?cmp</span><span class="w"> </span><span class="nn">wbt</span><span class="o">:</span><span class="ss">P5</span><span class="w"> </span><span class="nv">?wikidata</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">SERVICE</span><span class="w"> </span><span class="nn">&lt;https://query.wikidata.org/sparql&gt;</span><span class="w"> </span><span class="p">{</span><span class="w">
    </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q135222054</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P921</span><span class="w"> </span><span class="nv">?pfas</span><span class="w"> </span><span class="p">.</span><span class="w">
    </span><span class="nv">?pfas</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P31</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q113145171</span><span class="w"> </span><span class="p">.</span><span class="w">
    </span><span class="k">BIND</span><span class="w"> </span><span class="p">(</span><span class="nb">substr</span><span class="p">(</span><span class="nb">str</span><span class="p">(</span><span class="nv">?pfas</span><span class="p">),</span><span class="mi">32</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?wikidata</span><span class="p">)</span><span class="w">
  </span><span class="p">}</span><span class="w">
  </span><span class="k">BIND</span><span class="w"> </span><span class="p">(</span><span class="s2">"Q5145"</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?P21</span><span class="p">)</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>I actually had to add 5 PFAS compounds in the VHP4Safety Compound Wiki first. That follows the
<a href="https://chem-bla-ics.linkedchemistry.info/2016/03/20/adding-disclosures-to-wikidata-with.html">same procedure for how I have been adding chemical compounds to Wikidata</a>
(see also <a href="https://doi.org/10.26434/chemrxiv-2025-53n0w">this preprint</a>).
The input <code class="language-plaintext highlighter-rouge">cas.smi</code> has the (missing) SMILES, Wikidata QID, and English label:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>C(CS(=O)(=O)O)C(C(C(C(C(C(F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F       Q27063662       6:2 Fluorotelomer sulfonate
CN(CC(=O)O)S(=O)(=O)C(C(C(C(C(C(F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F Q126605979      MeFHxSAA
CN(CC(=O)O)S(=O)(=O)C(C(C(C(F)(F)F)(F)F)(F)F)(F)F       Q126682412      MeFBSAA
C(=O)(C(C(F)(F)F)(F)OC(C(C(F)(F)F)(F)F)(F)F)O[H]        Q29387971       2,3,3,3-tetrafluoro-2-(heptafluoropropoxy)propanoic acid
C(C(C(=O)O)(F)F)(OC(C(C(OC(F)(F)F)(F)F)(F)F)(F)F)F      Q81981675       4,8-Dioxa-3H-perfluorononanoic acid
</code></pre></div></div>

<p>For reference, this is the command line I used to create QuickStatement instructions:</p>

<div class="language-shell highlighter-rouge"><div class="highlight"><pre class="highlight"><code>groovy createWDitemsFromSMILES.groovy <span class="nt">-w</span> compoundcloud.wikibase.cloud <span class="nt">-c</span> Q2368 <span class="nt">-d</span> P5 <span class="nt">-l</span> <span class="nt">-i</span> wikidata <span class="nt">-a</span> P11
</code></pre></div></div>

<h2 id="final-remark">Final remark</h2>

<p>Are these 16 the only PFAS in our body? With 28 studied out of <a href="https://doi.org/10.1021/acs.est.3c04855">a potential seven million</a>,
I doubt it.</p>]]></content><author><name>Egon Willighagen</name></author><category term="pfas" /><category term="chemistry" /><category term="fair" /><category term="scholia" /><category term="wikidata" /><category term="vhp4safety" /><category term="doi:10.26434/CHEMRXIV-2025-53N0W" /><category term="cito:citesAsRecommendedReading:10.1021/acs.est.3c04855" /><summary type="html"><![CDATA[A recent report by the Dutch RIVM, PFAS in the blood of the Dutch population (doi:10.21945/RIVM-2025-0094), writes that seven PFAS compounds are found in blood samples of all tested people. Another nine compounds are found in at least 1-in-10 people. Because there is relevant data in the report on the 28 studied PFAS compound, I wanted to have the report more FAIR than it is on the website. Why this report? Well, the chemistry and the history is fascinating and brutal (I like this Veritasium video).]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pfas_report.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pfas_report.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">New preprint: “Scholia Chemistry: access to chemistry in Wikidata”</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata.html" rel="alternate" type="text/html" title="New preprint: “Scholia Chemistry: access to chemistry in Wikidata”" /><published>2025-05-25T00:00:00+00:00</published><updated>2025-05-25T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/05/25/new-preprint-scholia-chemistry-access-to-chemistry-in-wikidata.html"><![CDATA[<p>Two week ago I uploaded a paper that has been in the works for some time. In fact, I first mention it as conference paper
for the special issue of the <a href="https://scholia.toolforge.org/event/Q47501229">11th International Conference on Chemical Structures</a>,
you know, the meeting held in 2018, of which <a href="https://iccs-nl.org/">the 13th edition</a> starts in 7 days. I had a
<a href="https://doi.org/10.6084/m9.figshare.6356027.v1">poster</a> at that conference which I described in
<a href="https://chem-bla-ics.linkedchemistry.info/2018/08/18/compound-class-identifiers-in-wikidata.html">this blog post</a>.</p>

<p>In turn, that poster described work of at least three years, going back to
<a href="https://chem-bla-ics.linkedchemistry.info/2015/12/22/new-edition-getting-cas-registry.html">adding identifiers in 2015</a>
and <a href="https://chem-bla-ics.linkedchemistry.info/2016/01/27/adding-chemical-compound-to-wikidata.html">chemical structures in early 2016</a>.
I started <a href="https://chem-bla-ics.linkedchemistry.info/2016/03/20/adding-disclosures-to-wikidata-with.html">using scripts two months later</a>.
This helped a lot with <a href="https://chem-bla-ics.linkedchemistry.info/2016/03/27/migrating-pka-data-from-drugmet-to.html">migrating pKa data</a>
from a custom Semantic MediaWiki installation to Wikidata and with adding thousands of EPA CompTox
<a href="https://chem-bla-ics.blogspot.com/2017/01/epa-comptox-dashboard-ids-in-wikidata.html">identifiers in 2017</a>.</p>

<p>But that 2018 conference paper never happened. Because <a href="https://chem-bla-ics.linkedchemistry.info/2017/10/15/two-conference-proceedings.html">Scholia did</a>.
And even on the ICCS poster, Scholia was used to visualize chemistry data in Wikidata. To be honest, not just that,
of course. About a year ago I had a serious go at finishing the paper, and it was sent around to co-authors.
But I realized at the time, that the paper was lacking some good suggestions how the peer review our
actual contributions to Wikidata. I could hardly expect readers of the paper browse the individual
histories of all, by then, 1.3 million chemical compounds. And during the holidays I collected a few
tools, which I had lined up to add to the manuscript.</p>

<p>However, another thing happened, the COVID-19 pandemic. While all the experience helped a lot with getting
knowledge together around SARS-CoV-2, it also made something else clear: the software behind Wikidata
does not scale well (enough). This lead to plans to split the RDF graph representation into two
separate SPARQL endpoints. And that breaks many, if not most, of Scholia’s SPARQL queries, including
those for the chemistry aspects. The situation in Summer 2024 was that there was a significant
chance Scholia would not survive the split. And the <em>Scholia Chemistry</em> paper had to wait. You
cannot publish an article of which the website is gone before it is formally accepted.</p>

<p>Let me make clear, this graph split is not solved and the risk is not gone. But a serious of unfunded,
weekend hackathons allows us to refactor Scholia to give us a chance. It started with
<a href="https://chem-bla-ics.linkedchemistry.info/2024/08/23/scholia.html">making Scholia more configurable</a>.
We had the first hackathons in October and November, and I had
<a href="https://chem-bla-ics.linkedchemistry.info/2025/04/20/the-april-2025-scholia-hackathon.html">four more hackathon weekends</a>
this April.</p>

<p>The graph split into a main graph and a scholarly graph <a href="https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/WDQS_graph_split">happened on May 9</a>.
Currently, we have been granted extra time and can use a legacy server with the full graph, but a lot
less hardware, so slower. A final patch, merged in last week, allows us to define which SPARQL endpoint a query
should run. So, each time we port a SPARQL query, we can directly update Scholia, making the migration
somewhat more manageable.</p>

<p>But, with those uncertainties out of the way, it was time to finish the Scholia Chemistry paper!</p>

<p>The preprint (doi:<a href="https://doi.org/10.26434/chemrxiv-2025-53n0w">10.26434/chemrxiv-2025-53n0w</a>) brings
10 years of research together, and describes details of the used methods not formally peer-reviewed before.
We describe in detail how chemical structures are added, the choices of Wikidata on how to
represent chemical structures, how we curate the quality, and how we visualize chemical structures
and data with Scholia. As you can expect, the Chemistry Development Kit has an important role,
along with the InChI.</p>

<p>The paper introduces three new Scholia <em>aspects</em> for chemicals, chemical classes, and elements.
Each aspect is a template for a page with information about molecular entities and chemical substances,
compound classes (like <em>fatty acids</em>), and elements (like carbon). Each template provides relevant
information. Of course, any compound, class, or element can also still be opened in the Scholia
“topic” aspect, listing relevant literature.</p>

<p>With this paper we aim to show that Wikidata is a innovative platform that meets the needs for
a chemical structure database, with detailed data provenance, and scalable community curation.</p>

<p>I welcome your strongest peer review on the preprint. I don’t liking settling for anything less.
Here’s the abstract:</p>

<blockquote>
  <p>Sharing knowledge on chemicals in the digital age has been the playground of databases such
as the Chemical Abstract Services and PubChem. Wikipedia complements this field by providing
context to chemicals aimed at a broad audience, but is not easily read by machines. Wikidata
was started as a database service to improve the machine readability of the knowledge captured
in Wikipedia. Wikidata has an open license, application programming interfaces, and a strong
provenance model. Scholia uses the features to provide access to chemical knowledge. This
study reviews the chemistry in Wikidata, shows how thousands of new chemicals were added,
extends Wikidata with new properties for chemical representation and external links to
additional databases, and shows how we extended Scholia to represent the chemistry in Wikidata.</p>
</blockquote>

<p>Thanks to Finn, Denise, Daniel, and Adriano!</p>]]></content><author><name>Egon Willighagen</name></author><category term="wikidata" /><category term="scholia" /><category term="chemistry" /><category term="iccs" /><category term="cito:citesAsEvidence:10.6084/m9.figshare.6356027.v1" /><category term="doi:10.26434/CHEMRXIV-2025-53N0W" /><summary type="html"><![CDATA[Two week ago I uploaded a paper that has been in the works for some time. In fact, I first mention it as conference paper for the special issue of the 11th International Conference on Chemical Structures, you know, the meeting held in 2018, of which the 13th edition starts in 7 days. I had a poster at that conference which I described in this blog post.]]></summary></entry><entry><title type="html">Beilstein journals contain Bioschemas</title><link href="https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html" rel="alternate" type="text/html" title="Beilstein journals contain Bioschemas" /><published>2025-02-13T00:00:00+00:00</published><updated>2025-02-13T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html"><![CDATA[<p>Two weeks ago, the <a href="https://www.beilstein-journals.org/bjoc/news/LAFGBV6PT5ASC5R7JOKSEXOQYM">Beilstein Institute announced Bioschemas support in their journals</a>:</p>

<blockquote>
  <p>We streamline the discoverability of your research by incorporating machine-readable chemical information into many of our published articles.
This includes the conversion of chemical structures from submitted ChemDraw files to InChI strings and validating them using open-source tools.</p>
</blockquote>

<p>The idea is far from new and has been around for two decades. But the <a href="https://scholia.toolforge.org/publisher/Q4881267">two Beilstein journals</a>
(both <a href="https://en.wikipedia.org/wiki/Diamond_open_access">diamond Open Access</a>), actually integrated into their active publishing model.
That has been trialed and put in action before. For example, there was (is?) <a href="https://doi.org/10.59350/ne4rf-wey66">Project Prospect</a>
(2007), <a href="https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html">chemical structure annotation in Nature Chemistry</a>
(2009), <a href="https://chem-bla-ics.linkedchemistry.info/2014/02/21/slow-publishing-innovation.html">SMILES in the ACS Journal of Medicinal Chemistry</a>
(2014) (doi:<a href="https://doi.org/10.1021/jm5002056">10.1021/jm5002056</a>),
and <em>FAIR chemical structures in the Journal of Cheminformatics</em> (2021) (doi:<a href="https://doi.org/10.1186/s13321-021-00520-4">10.1186/s13321-021-00520-4</a>).</p>

<p>But this announcement is a new step. I like how validation of the chemical structures is part of the approach, and I like
how they use the <a href="https://bioschemas.org/">Bioschemas</a> extention of <a href="https://schema.org/">schema.org</a>. The last because
they use two Bioschemas types/profiles that contributed to or initiated, respectively: <a href="https://bioschemas.org/profiles/MolecularEntity/0.5-RELEASE">MolecularEntity</a>
and <a href="https://bioschemas.org/profiles/ChemicalSubstance/0.4-RELEASE">ChemicalSubstance</a>.</p>

<p>First stop for me is to check the schema.org annotation with a validation tool, like <a href="https://search.google.com/test/rich-results">Google’s Rich Results Test</a>.
That gives an idea how they may have have their search engine pick it up. The test article I was given on LinkedIn is
Xiao <em>et al.</em>’s <em>Molecular diversity of the reactions of MBH carbonates of isatins and various nucleophiles</em>
(doi:<a href="https://doi.org/10.3762/bjoc.21.21">10.3762/bjoc.21.21</a>) in the <a href="https://scholia.toolforge.org/venue/Q2894008">Beilstein Journal of Organic Chemistry</a>,
and we indeed <a href="https://search.google.com/test/rich-results/result?id=FRW9wBOpXtsMp9TLUV6SfQ">see the schema.org annotation show up</a>:</p>

<p><img src="/assets/images/bjoc_bioschemas.png" alt="" /></p>

<p>And because of the use of open standards, extracting the information is not so hard with, for example here,
Bacting (doi:<a href="https://doi.org/10.21105/joss.02558">10.21105/joss.02558</a>), based on a 2022 script from the NanoSafety Cluster
projects NanoCommons and SbD4Nano:</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'io.github.egonw.bacting'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'managers-rdf'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'1.0.4'</span><span class="o">)</span>
<span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'io.github.egonw.bacting'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'managers-ui'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'1.0.4'</span><span class="o">)</span>
<span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'io.github.egonw.bacting'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'net.bioclipse.managers.jsoup'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'1.0.4'</span><span class="o">)</span>

<span class="n">bioclipse</span> <span class="o">=</span> <span class="k">new</span> <span class="n">net</span><span class="o">.</span><span class="na">bioclipse</span><span class="o">.</span><span class="na">managers</span><span class="o">.</span><span class="na">BioclipseManager</span><span class="o">(</span><span class="s2">"."</span><span class="o">);</span>
<span class="n">rdf</span> <span class="o">=</span> <span class="k">new</span> <span class="n">net</span><span class="o">.</span><span class="na">bioclipse</span><span class="o">.</span><span class="na">managers</span><span class="o">.</span><span class="na">RDFManager</span><span class="o">(</span><span class="s2">"."</span><span class="o">);</span>
<span class="n">jsoup</span> <span class="o">=</span> <span class="k">new</span> <span class="n">net</span><span class="o">.</span><span class="na">bioclipse</span><span class="o">.</span><span class="na">managers</span><span class="o">.</span><span class="na">JSoupManager</span><span class="o">(</span><span class="s2">"."</span><span class="o">);</span>

<span class="n">articles</span> <span class="o">=</span> <span class="o">[</span>
   <span class="n">args</span><span class="o">[</span><span class="mi">0</span><span class="o">]</span>
<span class="o">]</span>

<span class="n">kg</span> <span class="o">=</span> <span class="n">rdf</span><span class="o">.</span><span class="na">createInMemoryStore</span><span class="o">()</span>

<span class="k">for</span> <span class="o">(</span><span class="n">article</span> <span class="k">in</span> <span class="n">articles</span><span class="o">)</span> <span class="o">{</span>
    <span class="n">htmlContent</span> <span class="o">=</span> <span class="n">bioclipse</span><span class="o">.</span><span class="na">download</span><span class="o">(</span><span class="n">article</span><span class="o">)</span>

    <span class="n">htmlDom</span> <span class="o">=</span> <span class="n">jsoup</span><span class="o">.</span><span class="na">parseString</span><span class="o">(</span><span class="n">htmlContent</span><span class="o">)</span>

    <span class="c1">// application/ld+json</span>

    <span class="n">bioschemasSections</span> <span class="o">=</span> <span class="n">jsoup</span><span class="o">.</span><span class="na">select</span><span class="o">(</span><span class="n">htmlDom</span><span class="o">,</span> <span class="s2">"script[type='application/ld+json']"</span><span class="o">);</span>

    <span class="k">for</span> <span class="o">(</span><span class="n">section</span> <span class="k">in</span> <span class="n">bioschemasSections</span><span class="o">)</span> <span class="o">{</span>
        <span class="n">bioschemasJSON</span> <span class="o">=</span> <span class="n">section</span><span class="o">.</span><span class="na">html</span><span class="o">()</span>
        <span class="n">rdf</span><span class="o">.</span><span class="na">importFromString</span><span class="o">(</span><span class="n">kg</span><span class="o">,</span> <span class="n">bioschemasJSON</span><span class="o">,</span> <span class="s2">"JSON-LD"</span><span class="o">)</span>
    <span class="o">}</span>
<span class="o">}</span>

<span class="n">turtle</span> <span class="o">=</span> <span class="n">rdf</span><span class="o">.</span><span class="na">asTurtle</span><span class="o">(</span><span class="n">kg</span><span class="o">);</span>

<span class="n">println</span> <span class="s2">"#"</span> <span class="o">+</span> <span class="n">rdf</span><span class="o">.</span><span class="na">size</span><span class="o">(</span><span class="n">kg</span><span class="o">)</span> <span class="o">+</span> <span class="s2">" triples detected in the JSON-LD"</span>
<span class="c1">// println turtle</span>


<span class="n">sparql</span> <span class="o">=</span> <span class="s2">"""
PREFIX schema: &lt;http://schema.org/&gt;
SELECT ?entity ?inchikey ?smiles WHERE {
  ?entity a schema:MolecularEntity .
  OPTIONAL { ?entity schema:inChIKey ?inchikey }
  OPTIONAL { ?entity schema:smiles ?smiles }
}
"""</span>

<span class="n">results</span> <span class="o">=</span> <span class="n">rdf</span><span class="o">.</span><span class="na">sparql</span><span class="o">(</span><span class="n">kg</span><span class="o">,</span> <span class="n">sparql</span><span class="o">)</span>

<span class="k">for</span> <span class="o">(</span><span class="n">i</span><span class="o">=</span><span class="mi">1</span><span class="o">;</span><span class="n">i</span><span class="o">&lt;=</span><span class="n">results</span><span class="o">.</span><span class="na">rowCount</span><span class="o">;</span><span class="n">i</span><span class="o">++)</span> <span class="o">{</span>
  <span class="n">println</span> <span class="s2">"${results.get(i, "</span><span class="n">inchikey</span><span class="s2">")}\t${results.get(i, "</span><span class="n">smiles</span><span class="s2">")}"</span>
<span class="o">}</span>
</code></pre></div></div>

<p>The output is a simple table:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>MGAPJMNPGGTFHJ-JEIPZWNWSA-N     CN1C(=O)/C(=C/2\C3=CC(=CC=C3N(CC4=CC=CC=C4)C2=O)Cl)/C(=P(C5=CC=CC=C5)(C6=CC=CC=C6)C7=CC=CC=C7)C1=O
XEWMQVUVGAHESA-UHFFFAOYSA-N     CC1=CC=C(C=C1)NC2=C(C3C4=CC(=CC=C4N(CC5=CC=CC=C5)C3=O)C)C(=O)N(C)C2=O
UVTJORFYHPGJDZ-PYCFMQQDSA-N     CCCCN1C2=CC=C(C)C=C2/C(=C(\C#N)/CNC3=CC=C(C)C=C3)/C1=O
ILWGDUYVQRAMMG-PGMHBOJBSA-N     CCCCN1C2=CC=C(C)C=C2/C(=C(\C#N)/CNC3=CC=C(C=C3)Cl)/C1=O
CAFIBKBZWJFZCW-FXBPSFAMSA-N     CCCCN1C2=CC=C(C)C=C2/C(=C(\C#N)/CNC3=CC=CC=C3)/C1=O
UOJSFLANMVIMBV-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=C(C=C4)Cl)C1=O
VNJBTGZXAGHCSO-OAPYJULQSA-N     COC(=O)/C(=C\1/C2=C(C=CC=C2)N(CC3=CC=CC=C3)C1=O)/C=P(C4=CC=CC=C4)(C5=CC=CC=C5)C6=CC=CC=C6
KJXQRAKSOANQTJ-GFMRDNFCSA-N     CC1=CC=C(C=C1)NC/C(=C\2/C3=C(C=CC=C3)N(CC4=CC=CC=C4)C2=O)/C#N
IGEBJMZDOPBFGF-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=CC=C4)C1=O
SSANVPNESOMKOM-AWQADKOQSA-N     C1=CC=C(C=C1)CN2C3=CC=C(C=C3/C(=C(/C#N)\C=P(C4=CC=CC=C4)(C5=CC=CC=C5)C6=CC=CC=C6)/C2=O)Cl
GEHWHSHQSIOZKL-NVQSTNCTSA-N     CCCCN1C2=CC=C(C=C2/C(=C\3/C(=P(C4=CC=CC=C4)(C5=CC=CC=C5)C6=CC=CC=C6)C(=O)N(C)C3=O)/C1=O)Cl
PALRSQOHFLRWDH-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=C(C=C4)OC)C1=O
KBFODZMDSAFLFR-UHFFFAOYSA-N     CN1C(=O)C(=C(C1=O)NC2=CC(=CC=C2)Cl)C3C4=CC(=CC=C4N(CC5=CC=CC=C5)C3=O)Cl
JCGAVVZYXDJPBU-GFMRDNFCSA-N     CC1=C(C=CC=C1)NC/C(=C\2/C3=C(C=CC=C3)N(CC4=CC=CC=C4)C2=O)/C#N
DZFPCPDEQGLPLY-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=C(C)C=C4)C1=O
XMRNJCJUOXYXJU-DAFNUICNSA-N     CC1=CC=C(C=C1)NC/C(=C\2/C3=CC(=CC=C3N(CC4=CC=CC=C4)C2=O)C)/C#N
SSDSNBBHEUUKGI-UHFFFAOYSA-N     CC1=CC=C2C(=C1)C(C3=C(C(=O)N(C)C3=O)N(C)C4=CC=CC=C4)C(=O)N2CC5=CC=CC=C5
USFYPRDMNXMWPO-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=C(C=C4)Br)C1=O
XYHTWFULRHTEAG-MUGXBBEHSA-N     CCCCN1C2=CC=C(C)C=C2/C(=C(/C#N)\C=P(C3=CC=CC=C3)(C4=CC=CC=C4)C5=CC=CC=C5)/C1=O
XALDZIBHNNIVAM-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NC4=C(C=CC=C4)O)C1=O
TUTWQHBRQPMLME-OAPYJULQSA-N     COC(=O)/C(=C\1/C2=CC(=CC=C2N(CC3=CC=CC=C3)C1=O)Cl)/C=P(C4=CC=CC=C4)(C5=CC=CC=C5)C6=CC=CC=C6
IYEHFTMZZMIPRU-UHFFFAOYSA-N     CC1=CC=C(C=C1)NC2=C(C3C4=CC(=CC=C4N(CC5=CC=CC=C5)C3=O)Cl)C(=O)N(C)C2=O
KBSDGNPLIPXCEX-UHFFFAOYSA-N     CCCCN1C2=CC=C(C)C=C2C(C3=C(C(=O)N(C)C3=O)NCC4=CC=CC=C4)C1=O
BQGIUMITIGHBSD-UHFFFAOYSA-N     CCCCNC1=C(C2C3=CC(=CC=C3N(CC4=CC=CC=C4)C2=O)C)C(=O)N(C)C1=O
PNSOLOPHIVUPOZ-MNDPQUGUSA-N     CCCCNC/C(=C\1/C2=CC(=CC=C2N(CCCC)C1=O)C)/C#N
HLTBKJRJOIZCMJ-PYCFMQQDSA-N     CCCCN1C2=CC=C(C)C=C2/C(=C(\C#N)/CN(C)C3=CC=CC=C3)/C1=O
FFLHFLUBMRBQTB-UHFFFAOYSA-N     CCCCN1C2=CC=C(C=C2C(C3=C(C(=O)N(C)C3=O)NC4=CC=C(C)C=C4)C1=O)F
FOQOVOLYYARWPA-NKFKGCMQSA-N     C1=CC=C(C=C1)CN2C3=C(C=CC=C3)/C(=C(\C#N)/CNC4=CC(=CC=C4)Cl)/C2=O
KLEPCAQFOXJLNV-UHFFFAOYSA-N     CC1=C(C=CC=C1)NC2=C(C3C4=CC(=CC=C4N(CC5=CC=CC=C5)C3=O)Cl)C(=O)N(C)C2=O
</code></pre></div></div>

<p>That also made me realize that there are not chemical names in the annotation. That would be really useful to move things
forward. Then again, PubChem will likely just generate the IUPAC name, since they have access to such software anyway.
They have teamed up with PubChem which will index it, but I will be interested in seeing how to use this for
<code class="language-plaintext highlighter-rouge">main subject</code> annotation in <a href="https://www.wikidata.org/wiki/Wikidata:WikiProject_Chemistry">Wikidata</a>.</p>

<p>A final note for now, the model they use is annotate the article with chemical substances (<code class="language-plaintext highlighter-rouge">ChemicalSubstance</code>) with
(one or more?) molecular entities (`MolecularEntity’). That is a model that scales well to their other journal,
the <a href="https://scholia.toolforge.org/venue/Q814756">Beilstein Journal of Nanotechnology</a>. But scraping that is for another post.</p>]]></content><author><name>Egon Willighagen</name></author><category term="bioschemas" /><category term="rdf" /><category term="chemistry" /><category term="cito:citesForInformation:10.59350/ne4rf-wey66" /><category term="cito:citesForInformation:10.1186/s13321-021-00520-4" /><category term="cito:citesForInformation:10.1021/jm5002056" /><category term="cito:usesDataFrom:10.3762/bjoc.21.21" /><category term="cito:usesMethodIn:10.21105/joss.02558" /><category term="cito:citesForInformation:10.59350/40377-hz881" /><category term="beilstein" /><summary type="html"><![CDATA[Two weeks ago, the Beilstein Institute announced Bioschemas support in their journals:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/bjoc_bioschemas.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/bjoc_bioschemas.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Two meetings: ELIXIR Toxicology and FAIR4ChemNL</title><link href="https://chem-bla-ics.linkedchemistry.info/2024/06/10/two-meetings.html" rel="alternate" type="text/html" title="Two meetings: ELIXIR Toxicology and FAIR4ChemNL" /><published>2024-06-10T00:00:00+00:00</published><updated>2024-06-10T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2024/06/10/two-meetings</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2024/06/10/two-meetings.html"><![CDATA[<p>Noting that in the coming week I am not attending the <a href="https://elixir-europe.org/events/elixir-all-hands-2024">ELIXIR All Hands in Uppsala</a>.
Having lived in (and around) Uppsala for more than three years, I am disappointed and with the first stories from colleagues coming
in even more. But it has been a way too busy year, I have much to finish up, and I need to take care of myself too. I am not 32 anymore.</p>

<p>But in the past two weeks I did attend two workshops. The first was a <a href="https://www.aanmelder.nl/intoxicom2024firstworkshop">workshop</a> by the
<a href="https://elixir-europe.org/communities/toxicology">ELIXIR Toxicology Community</a>, which was held in Utrecht/NL. The programme was around
FAIR and included two really nice hands-on sessions where we developed drafts for <a href="https://faircookbook.elixir-europe.org/">FAIR Cookbook</a>
recipes (see also doi:<a href="https://doi.org/10.1038/s41597-023-02166-3">10.1038/s41597-023-02166-3</a>) and for
<a href="https://www.go-fair.org/how-to-go-fair/fair-implementation-profile/">FAIR Implementation Profiles</a>
(doi:<a href="https://doi.org/10.1007/978-3-030-65847-2_13">10.1007/978-3-030-65847-2_13</a>). We will write up a
<a href="https://biohackrxiv.org/discover">BioHackrXiv</a> report.</p>

<p>The second workshop was last week, the <a href="https://tdcc.nl/evenementen/fair4chemnl-workshop/">FAIR4ChemNL workshop</a>, which was also held
in Utrecht/NL. The topic was FAIR in chemistry, and we discussed various aspects. There was a significant participant group from the
German NFDI4Cat project (“Cat” is short for (chemical) catalysis), which recently published a nice analysis of several ontologies
(doi:<a href="https://doi.org/10.1186/s13321-024-00807-2">10.1186/s13321-024-00807-2</a>). And there was also a lot of mention of RDF and SPARQL.</p>

<p>I think it is time for a new special issue around semantic web technologies.</p>]]></content><author><name>Egon Willighagen</name></author><category term="elixir" /><category term="fair" /><category term="chemistry" /><category term="doi:10.1038/S41597-023-02166-3" /><category term="justdoi:10.1007/978-3-030-65847-2_13" /><category term="justdoi:10.1186/S13321-024-00807-2" /><category term="rdf" /><category term="sparql" /><category term="fair4chemnl" /><summary type="html"><![CDATA[Noting that in the coming week I am not attending the ELIXIR All Hands in Uppsala. Having lived in (and around) Uppsala for more than three years, I am disappointed and with the first stories from colleagues coming in even more. But it has been a way too busy year, I have much to finish up, and I need to take care of myself too. I am not 32 anymore.]]></summary></entry><entry><title type="html">Boiling points in Wikidata</title><link href="https://chem-bla-ics.linkedchemistry.info/2023/08/12/boiling-points-in-wikidata.html" rel="alternate" type="text/html" title="Boiling points in Wikidata" /><published>2023-08-12T00:00:00+00:00</published><updated>2023-08-12T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2023/08/12/boiling-points-in-wikidata</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2023/08/12/boiling-points-in-wikidata.html"><![CDATA[<p>Some days ago, I started added boiling points to <a href="https://wikidata.org/">Wikidata</a>, referenced from
<a href="https://scholia.toolforge.org/work/Q22236188">Basic Laboratory and Industrial Chemicals</a> (wikidata:Q22236188),
<a href="https://scholia.toolforge.org/author/Q18609741">David R. Lide</a>’s
‘a CRC quick reference handbook’ from 1993 (well, the edition I have). But Wikidata
<a href="https://www.wikidata.org/wiki/User_talk:Egon_Willighagen#Basic_laboratory_and_industrial_chemicals:_a_CRC_quick_reference_handbook_(Q22236188)">wants</a>
pressure (wikidata:P2077) info at which the boiling point (wikidata:P2102) was measured. Rightfully so. But I had not added those yet,
because it slows me and can be automated with <a href="https://quickstatements.toolforge.org/">QuickStatements</a>.</p>

<p>I just need a few SPARQL queries to list to which statements the qualifiers needs to be added. Basically, all boiling points which has the
book as a reference and that do not have the pressure info. First, there are values with ‘unknown value’, which results in blank nodes
(by the time you read this, they likely are already fixed):</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span><span class="w"> </span><span class="nv">?cmp</span><span class="w"> </span><span class="nv">?bp</span><span class="w"> </span><span class="nv">?pressure</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?cmp</span><span class="w"> </span><span class="nn">p</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">prov</span><span class="o">:</span><span class="ss">wasDerivedFrom</span><span class="o">/</span><span class="nn">pr</span><span class="o">:</span><span class="ss">P248</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q22236188</span><span class="w"> </span><span class="p">;</span><span class="w">
    </span><span class="nn">ps</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?bp</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">pq</span><span class="o">:</span><span class="ss">P2077</span><span class="w"> </span><span class="nv">?pressure</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">FILTER</span><span class="w"> </span><span class="p">(</span><span class="nb">contains</span><span class="p">(</span><span class="nb">str</span><span class="p">(</span><span class="nv">?pressure</span><span class="p">),</span><span class="w"> </span><span class="s2">"http://"</span><span class="p">))</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>So, to get the list for which I want to write the QuickStatements which does not have any P2077 qualifier yet, I use
<a href="https://query.wikidata.org/#SELECT%20%3Fcmp%20WHERE%20%7B%0A%20%20%3Fcmp%20p%3AP2102%20%3FbpStatement%20.%0A%20%20%3FbpStatement%20prov%3AwasDerivedFrom%2Fpr%3AP248%20wd%3AQ22236188%20%3B%0A%20%20%20%20ps%3AP2102%20%3Fbp%20.%0A%20%20MINUS%20%7B%20%3FbpStatement%20pq%3AP2077%20%3Fpressure%20%7D%0A%7D">this query</a>:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span><span class="w"> </span><span class="nv">?cmp</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?cmp</span><span class="w"> </span><span class="nn">p</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">prov</span><span class="o">:</span><span class="ss">wasDerivedFrom</span><span class="o">/</span><span class="nn">pr</span><span class="o">:</span><span class="ss">P248</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q22236188</span><span class="w"> </span><span class="p">;</span><span class="w">
    </span><span class="nn">ps</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?bp</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">MINUS</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">pq</span><span class="o">:</span><span class="ss">P2077</span><span class="w"> </span><span class="nv">?pressure</span><span class="w"> </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>At the time of writing, this lists 54 boiling points.</p>

<p>I can the WDQS create CSV-styled QuickStatements with:</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">SELECT</span><span class="w"> </span><span class="p">(</span><span class="nb">SUBSTR</span><span class="p">(</span><span class="nb">STR</span><span class="p">(</span><span class="nv">?cmp</span><span class="p">),</span><span class="mi">32</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?qid</span><span class="p">)</span><span class="w"> </span><span class="nv">?P2102</span><span class="w"> </span><span class="nv">?qal2077</span><span class="w"> </span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?cmp</span><span class="w"> </span><span class="nn">p</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">prov</span><span class="o">:</span><span class="ss">wasDerivedFrom</span><span class="o">/</span><span class="nn">pr</span><span class="o">:</span><span class="ss">P248</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q22236188</span><span class="w"> </span><span class="p">;</span><span class="w">
    </span><span class="nn">ps</span><span class="o">:</span><span class="ss">P2102</span><span class="w"> </span><span class="nv">?P2102</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">MINUS</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?bpStatement</span><span class="w"> </span><span class="nn">pq</span><span class="o">:</span><span class="ss">P2077</span><span class="w"> </span><span class="nv">?pressure</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">BIND</span><span class="w"> </span><span class="p">(</span><span class="s2">"101.325U21064807"</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?qal2077</span><span class="p">)</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>Here, the SPARQL variables double as QuickStatement instructions. Finally, note to use of “U21064807” which is the Wikidata item for
kilopascal (wikidata:Q21064807).</p>

<p>I also need to “add” the boiling point again, to make sure QuickStatements knows which statement to add the qualifier to. I think this
can be done better, but not sure how to target statements directly. This is not fool proof: I noted that this approach ignores the
situation where there are two statements with the (exact) same boiling point, but different error margins. But that I will monitor
and where needed correct manually.</p>]]></content><author><name>Egon Willighagen</name></author><category term="rdf" /><category term="wikidata" /><category term="chemistry" /><summary type="html"><![CDATA[Some days ago, I started added boiling points to Wikidata, referenced from Basic Laboratory and Industrial Chemicals (wikidata:Q22236188), David R. Lide’s ‘a CRC quick reference handbook’ from 1993 (well, the edition I have). But Wikidata wants pressure (wikidata:P2077) info at which the boiling point (wikidata:P2102) was measured. Rightfully so. But I had not added those yet, because it slows me and can be automated with QuickStatements.]]></summary></entry><entry><title type="html">new: “CAS Common Chemistry in 2021: Expanding Access to Trusted Chemical Information for the Scientific Community”</title><link href="https://chem-bla-ics.linkedchemistry.info/2022/05/22/new-cas-common-chemistry-in-2021.html" rel="alternate" type="text/html" title="new: “CAS Common Chemistry in 2021: Expanding Access to Trusted Chemical Information for the Scientific Community”" /><published>2022-05-22T00:00:00+00:00</published><updated>2022-05-22T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2022/05/22/new-cas-common-chemistry-in-2021</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2022/05/22/new-cas-common-chemistry-in-2021.html"><![CDATA[<p>Open Science is happening. The merits are no longer theoretical or idealistic but tangible. Research is faster than ever, more vetted than ever (think PubPeer),
more cited than ever. Fairly, not just because of Open Science, but open access causes readership causes impact causes citations. When new people and
organizations start adopting Open Science this warms my hearth.</p>

<p>So, when I was asked to work with <a href="https://www.cas.org/">Chemical Abstracts Service</a> (CAS) on a new, bigger than ever version of
<a href="https://commonchemistry.cas.org/">Common Chemistry</a> (which started as a project between CAS and Wikipedia), I welcomed the project. I don’t quite
remember the first meetings, but roughly my task became to work with the new content and match this against Wikidata and Wikipedia. It aligned well
with <a href="https://bridgedb.github.io/">BridgeDb</a>, <a href="https://scholia.toolforge.org/">Scholia</a>, and our metabolomics research, so I even could find
sufficient research time for it. This work is now published in the <a href="https://pubs.acs.org/journal/jcisd8">JCIM</a>:
<em>CAS Common Chemistry in 2021: Expanding Access to Trusted Chemical Information for the Scientific Community</em>
(doi:<a href="https://doi.org/10.1021/acs.jcim.2c00268">10.1021/acs.jcim.2c00268</a>).</p>

<p><img src="/assets/images/images_medium_ci2c00268_0003.png" alt="" /> <br />
<em>Figure 2 from the article. Detailed record for caffeine in CAS Common Chemistry (image: CC-BY).</em></p>

<p>About Wikidata, the paper writes (CC-BY):</p>

<blockquote>
  <p>The latest release of CAS Common Chemistry has also supported updates and corrections to CAS RNs in Wikidata and Wikipedia. (22)
InChIKeys were calculated from CAS SMILES using Bacting 0.0.31 (23) with the Chemistry Development Kit 2.7.1 (24) and were
matched with content in Wikidata. The CAS RNs were then compared. References to CAS Common Chemistry were added for CAS RNs
that matched. Mismatches have been shared with the Wikidata and Wikipedia communities so that they can manually review and
correct the misleading entries using CAS Common Chemistry as a reference. Because Wikidata also curates identifiers from
other data sources, validated CAS RNs in Wikidata may also be used to cross-reference with other resources. Scripts are
provided in the Supporting Information.</p>
</blockquote>

<p>The alignment is a continuous process, as new chemical compounds get added to Wikidata on a weekly basis. The comparison of
Common Chemistry with Wikidata and Wikipedia resulted in a wealth of curation data, e.g. inconsistent CAS numbers linked to
InChIKeys, where Common Chemistry had a different match than Wikidata or Wikipedia.</p>

<p>CAS registry numbers were not added to Wikidata in this process, only confirmed or reported as different. The latter
allowed manual curation by the community, which it did. Reports <a href="https://www.wikidata.org/wiki/Wikidata_talk:WikiProject_Chemistry/CAS_Validation_Results">look like this</a>.
When a InChIKey-CAS RN combination in Wikidata was confirmed, it was recorded as a reference, like this:</p>

<p><img src="/assets/images/Screenshot_20220522_084233.png" alt="" /> <br />
<em>Screenshot of Wikidata with two references, one reflecting a confirmation
by the English Wikipedia (potentially the result of the original Common Chemistry
project) and the second as outcome of the now published project.</em></p>

<p>Thanks to everyone on this project and <a href="https://orcid.org/0000-0001-9316-9400">Andrea Jacobs</a>
particularly for leading this open science project.</p>]]></content><author><name>Egon Willighagen</name></author><category term="curation" /><category term="chemistry" /><category term="cas" /><category term="doi:10.1021/ACS.JCIM.2C00268" /><category term="bioclipse" /><category term="cdk" /><category term="bridgedb" /><summary type="html"><![CDATA[Open Science is happening. The merits are no longer theoretical or idealistic but tangible. Research is faster than ever, more vetted than ever (think PubPeer), more cited than ever. Fairly, not just because of Open Science, but open access causes readership causes impact causes citations. When new people and organizations start adopting Open Science this warms my hearth.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/images_medium_ci2c00268_0003.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/images_medium_ci2c00268_0003.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">ChemCuration 2019 Poster Conference: Call for Posters</title><link href="https://chem-bla-ics.linkedchemistry.info/2019/10/14/chemcuration-2019-poster-conference.html" rel="alternate" type="text/html" title="ChemCuration 2019 Poster Conference: Call for Posters" /><published>2019-10-14T00:00:00+00:00</published><updated>2019-10-14T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2019/10/14/chemcuration-2019-poster-conference</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2019/10/14/chemcuration-2019-poster-conference.html"><![CDATA[<p><span style="width: 40%; display: block; margin-left: auto; margin-right: auto; float: right">
<img src="/assets/images/Screenshot_20191014_174204.png" /> <br />
Twitter profile.
</span></p>

<p><em>It giet oan!</em> That it a Frisian phrase for something unlike is going to happen, like and particularly related to the
<a href="https://en.wikipedia.org/wiki/Elfstedentocht">Elfstedentocht</a>.</p>

<p><strong>ChemCuration 2019</strong> is a go. The <a href="https://chemcuration.github.io/chemcuration2019/">website is online</a>, the
<a href="https://twitter.com/chemcuration">Twitter account</a> and <a href="https://twitter.com/hashtag/chemcur2019">hashtag are ready</a>,
we got a poster prize, and here is the call for posters!</p>

<blockquote>
  <p>On December 3 the first ChemCuration conference will take place. ChemCuration 2019 is a one day, online-only conference around data curation and curated data in the chemistry domain. During the entire conference day, you can participate by tweeting about the poster that you uploaded, along with the meeting hashtag, and responding to questions about your poster in the 24 hours of the conference day. The poster must be available in an online repository (e.g. Zenodo or Figshare) under the CCZero, CC-BY or CC-BY-SA license prior to the conference.</p>

  <p>This is the meeting scope: anything around data curation and curated data of open science data in chemistry. This includes but is not limited to: 1. a new release of curated open data; 2. FAIR metadata around open data; and 3. open source tools for data curation.</p>

  <p><strong>How do I participate in ChemCuration?</strong><br />
You can participate in this online poster conference by presenting your poster on Twitter
during the conference day. You do this by first archiving your poster via Figshare or Zenodo,
with an open license (e.g. CCZero or CC-BY). Then, during the day you tweet an image of
(part of) your digital poster with the <a href="https://twitter.com/hashtag/chemcur2019">#chemcur2019</a>
hashtag, a short summary, and a link to your online poster with its DOI. The archived poster
should be a regular A0 poster (WxH = 841 x 1189 mm or 33.1 x 46.8 in)</p>

  <p><strong>Do I need to register?</strong><br />
Registration is not obligatory to participate. However, if you would like to be eligible
for a poster prize, then registration is required, by Nov. 30th, 2019. The registration form
is found at <a href="https://github.com/chemcuration/chemcuration2019/issues/new/choose">https://github.com/chemcuration/chemcuration2019/issues/new/choose</a></p>

  <p>More information can be found on the website (<a href="https://chemcuration.github.io/chemcuration2019/">https://chemcuration.github.io/chemcuration2019/</a>)
and on Twitter <a href="https://twitter.com/chemcuration">https://twitter.com/chemcuration</a></p>
</blockquote>]]></content><author><name>Egon Willighagen</name></author><category term="curation" /><category term="chemistry" /><summary type="html"><![CDATA[Twitter profile.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/Screenshot_20191014_174204.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/Screenshot_20191014_174204.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">ChemCuration: a small trick to fix the SMILES of glucuronides</title><link href="https://chem-bla-ics.linkedchemistry.info/2019/10/09/chemcuration-small-trick-to-fix-smiles.html" rel="alternate" type="text/html" title="ChemCuration: a small trick to fix the SMILES of glucuronides" /><published>2019-10-09T00:00:00+00:00</published><updated>2019-10-09T00:00:00+00:00</updated><id>https://chem-bla-ics.linkedchemistry.info/2019/10/09/chemcuration-small-trick-to-fix-smiles</id><content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2019/10/09/chemcuration-small-trick-to-fix-smiles.html"><![CDATA[<p><span style="width: 40%; display: block; margin-left: auto; margin-right: auto; float: right">
<img src="/assets/images/Screenshot_20191008_144049.png" /> <br />
Glucuronide functional group.
</span></p>

<p>Now that the <a href="https://chemcuration.github.io/chemcuration2019/">ChemCuration 2019</a> online poster conference is nearing, and
my upcoming talks about chemistry in <a href="https://wikidata.org/">Wikidata</a> (also needing curation), and the much longer process
of curation of metabolite (-like) structures in <a href="https://wikipathways.org/">WikiPathways</a>, I decided that something I
tweeted earlier this week is actually quite useful, and therefore something I should really write up in my lab notebook.</p>

<p><a href="https://en.wikipedia.org/wiki/Glucuronide">Glucuronide</a> is an example (biological) functional group. And there are several
databases that represent the stereochemistry now always correct. That is an interoperability (and thus FAIR) problem.
Correcting this is not trivial, particularly if you have to redraw the same glucuronide group again and again.</p>

<p>So, not looking forward to that, I invested a bit of time to find a <a href="http://opensmiles.org/">SMILES</a> trick. What if I had
a SMILES snippet that I could easily copy/paste and attach to the SMILES of the chemical structure it is attached to? Here
goes.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>O1[C@H](C(O)=O)[C@H]([C@H](O)[C@@H](O)[C@@H]1O9)O.
</code></pre></div></div>

<p>I just realized that <a href="https://twitter.com/egonwillighagen/status/1181573810543321088">the original 3 I used</a> can better be
a <code class="language-plaintext highlighter-rouge">9</code>, which is less likely to occur in the SMILES of the rest of the molecule. The period at the end is also deliberate.
That way, I can just copy past the SMILES of the rest directly after that period. Then I get a disconnected structure, but
I only have to put a 9 next to the atom that is binding to the glucuronide. So, let’s see the R group is methane, I get:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>O1[C@H](C(O)=O)[C@H]([C@H](O)[C@@H](O)[C@@H]1O9)O.C9
</code></pre></div></div>

<p>Now, next stop: <code class="language-plaintext highlighter-rouge">CoA</code> and other common biological tags.</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemistry" /><category term="curation" /><category term="wikidata" /><category term="wikipathways" /><category term="smiles" /><summary type="html"><![CDATA[Glucuronide functional group.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/Screenshot_20191008_144049.png" /><media:content medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/Screenshot_20191008_144049.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>