<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator>
  <link href="https://chem-bla-ics.linkedchemistry.info/feed.xml" rel="self" type="application/atom+xml"/>
  <link href="https://chem-bla-ics.linkedchemistry.info" rel="alternate" type="text/html"/>
  <updated>2026-08-14T05:35:03+00:00</updated>
  <id>https://chem-bla-ics.linkedchemistry.info/archive.xml</id>
  <title type="html">chem-bla-ics</title>
  <subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle>
  <author>
    <name>Egon Willighagen</name>
    <uri>https://orcid.org/0000-0001-7542-0286</uri>
  </author>

  
  <entry>
    <title type="html">Scientias, oudheidkunde, en wetenschap, hoe dan</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/08/07/scientias-oudheidkunde-wetenschap-hoe-dan.html" rel="alternate" type="text/html" title="Scientias, oudheidkunde, en wetenschap, hoe dan"/>
    <published>2026-08-07T00:00:00+00:00</published>
    <updated>2026-08-14T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/jtx6g-4d547</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/08/07/scientias-oudheidkunde-wetenschap-hoe-dan.html">
      <![CDATA[ <p>Communication of science is important to me and it comes in many formats. Educational resources, research articles, text books, popular science magazines, blogs,
TV shows, podcasts, etc (in no particular order). Somewhere during the pandemic, I started listening to podcasts. I had done that, in the early days,
but those early podcasts… well, let’s say, the format had not materialized yet. But also that it can be quite helpful to not play the podcast at the
normal speeds, but at, say 1.25x. Podcasts, like any of the above format have a style, an audience in mind. Maybe the audience is first year students,
peer in the research field, or the general audience. Or a politician that wants to get credibility. There are several podcasts that I have been listening.
Not as education. The may have episode notes, but not proper scientific references. References is one thing important in the communication. Context is another.</p>

<p>I have many books in my house on these aspects of science. From text books, to primary literature, to links to blog posts. It is not my research field,
and (scientific) communication is very much a research field. “It makes sense”, does not do, but the awareness of some of the things should be basic
academic training. Or probably earlier. Every writing has context. Not a random pointer, but a friend has recently written with several others an
interesting chapter on reading scientific literature (doi:<a href="https://doi.org/10.1515/9783110782844-010">10.1515/9783110782844-010</a>,
<a href="https://uplopen.com/reader/chapters/pdf/10.1515/9783110782844-010">pdf</a>).</p>

<h2 id="scientias">Scientias</h2>

<p>And while I started the <a href="https://scientias.nl/nieuws/video/podcast/">Scientias</a> (with (<em>the</em>) <a href="https://mastodon.nl/@Diederikjekel">Diederik Jekel</a> and
<a href="https://mastodon.nl/@krijnsoeteman">Krijn Soeteman</a>) episode to hear more about
the science behind “history”, there were several topics in the [Over oudheidkunde, nepvondsten en AI: wat weten we écht over het verleden?] episode
about the above things, with <a href="https://mastodon.social/@JonaLendering">Jona Lendering</a> as guest.</p>

<p><img src="/assets/images/scientias_episode.png" alt="" /></p>

<p>The things I heard that I got excited about (enough to sit down and write down these thougths) were small comments, other things were
more prominent. One more obvious discussion was that a lot of science news is focusing on new facts, and not the science behind it. Not just
the methods, but also the stack of assumptions, how they were combined,
valued, etc. As Lendering explains how a researcher does that (translated into English by me, with assumptions on the true intention):</p>

<blockquote>
  <p>This is a hypothesis. Now I stack this hypothesis on the previous hypothesis, and I stack another hypothesis on that, and then I stack
a hypothesis on that.</p>
</blockquote>

<p>Most science news outlets find that too complicated (for their audience). To me it feels they think their audience is too stupid to understand
the full story. Now, the Scientias Podcast made it a unique selling point to explain how those facts were derived from experiments, because they
find the full background and experiments behind our knowledge important, exciting. and relevant. I do.</p>

<p>And if I understood Lendering correctly, he too says there actually is plenty of room of that. The full story is more complicated and I cannot
do this justice. There is a growing body of scientific literature that cannot be captured in a single blog post. I will not even try. It will
likely just <a href="https://mainzerbeobachter.com/2017/07/31/mom-het-backfire-effect/">backfire</a>.</p>

<p>And I agree. That is not relevant to them nor their point, but it does explain I was excited about the post. It is always nice to hear like-minded people.
I do not hear this position about scientific communication a lot, but very, very much agree and try to practice.</p>

<h2 id="stacking">Stacking</h2>

<p>Lendering about the book he described in the above quote:</p>

<blockquote>
  <p>I think the conclusions are absolutely wrong, because it stacked too many uncertainties in a row.</p>
</blockquote>

<p>I love this quote. Lendering is talking here about archeology, but it applies to at least chemistry and the life sciences too.
I once wrote a grant application to the Swedisch national research funder VR, their equivalent of the NWO
on that topic (which got rejected, <em>because the candidate has too many international collaborations</em>,
despite scores in the fundable range).</p>

<p>It also very much resonates with the whole topic of LLMs, where exactly that is happening: the stacking of many uncertainties.
The real problem of AI since the 2000s is not the availability of great machine learning methods, but the facts to learn from,
and how well those methods can put those facts in perspective (think Red Riding Hood). That is why my research focus swifted
from chemometrics (the old name for <em>AI in chemistry</em>) and knowledge representation, to knowledge representation and interoperability.</p>

<h2 id="the-science-behind-archeology-and-history">The science behind archeology and history</h2>

<p>Most of the podcast is about archeology and how history is discovered. And the whole podcast does a wonderful job at explaining
how much has changed in the past 30 years. What I learned as a kid about the Romans has been updated in many ways. The the oldest
known documented genocide, in Belgica that started with a Roman military camp on the hill that I can see from my work office, has
peaked my interest in Roman history. But what fascinates me most are not the narratives but the science behind it. The new
chemical and biological approaches giving new evidence that puts archeological evidence in a richer context. Just listen (or
watch) the podcast. And I haven’t read it, but I understood <a href="https://mainzerbeobachter.com/mijn-boeken/oudheidkunde-is-een-wetenschap/">Oudheidkunde is een wetenschap</a>
covers a good bit of it.</p>

<p>Do I think that Nijmegen is the oldest Dutch city, and not Maastricht? Of course, I lived in Nijmegen for a large part of my life.
The competition between the place I studied before and study now is just great amusement. Bring on the popcorn.
And that aquaduct in Nijmegen? I cycled everyday on a Roman dike used for that aquaduct. As a researcher, I really
do not care the exact dates, or what material that aquaduct was made of. How we could learn those things, well,
yeah, the chemistry of physics behind that is cool too. Bring on the nerdy science.</p>

<h2 id="what-science-should-be-like">What science should be like</h2>

<p>Another topic in the above theme is what science should be. Lendering is negative about the research schools and
research institutes (in the Dutch implementation) and explains that different disciplines should communicate and collaborate
more. Since my field is in between two established research fields, I can painfully confirm that many (Dutch) researchers
are much to focused on their own narrow specialism, that they lost context to put that specialism in a wider context.</p>

<blockquote>
  <p>You only start writing a publication, after you collected and looked at all evidence.</p>
</blockquote>

<p>Lendering then talks about how modern science may have deviated too much from this. For people who have been following
the Open Science news, they have seen a good body of scientific literature that show that this indeed is not always
to case (sic). For example, Lendering states:</p>

<blockquote>
  <p>There is no substantial peer review anymore. Really weird publications happen, that should never have been published.</p>
</blockquote>

<p>Really, bold, perhaps, but sadly too close to reality with enough primary literature that has studied these issues
(some of which you can find convered in my blog posts).
Commercial publishers expect a peer review in 10 working days. They value speed over quality. It is there business model.
But also researchers that simply submit the manuscript as-is to the next journal, when the previous journal clearly outlined
limitations. There are reasons why I had to step down as editor from Springer Nature.</p>

<h2 id="calls-for-action">Calls for Action</h2>

<p>Before I write a blog post that takes you more time than to listen to the podcast, there are two more quotes I like to cover:</p>

<blockquote>
  <p>I could not resist the opportunity to mention paywalls. And what has been behind paywalls is not released
as open access. And there are not plans for this.</p>
</blockquote>

<p>Lendering continues with a good example from desinformation on why this must change. Again, I am very excited about
this statement, which I have been arguing for too.</p>

<p>The next quote I shortened a bit, because I want you to listen to the full discussion in the podcast, but here
I do not want to focus on the very convincingly bad urgency of the given example:</p>

<blockquote>
  <p>there is a theory from the 19th century, and that theory has been disproved since then. [..]
But due to digitization projects, old literature is available, with the old disproven theories too.
This is in itself not an issue. However, because the correcting literature from the last 30 years
are hidden behind the paywall, is the disproven theory unchallenged.</p>
</blockquote>

<p>There are many aspects of just this part, and many relate to how knowledge and how we got to knowledge
is spread, reused, etc. Bascially, the FAIR data principles are in that sense just a reformalisation of a much
older and bigger problem. And I think Lendering is spot on with this observation. Translating this to the
bigger issues, the commercial publishers can lecture researchers about FAIR data, but as long as the
with strong determination frustrate scientific progress with paywalls, etc, then we will not make the
progress we deserve and need.</p>

<h3 id="action-1-all-dutch-literature-must-become-green-open-access">Action 1: all Dutch literature must become green Open Access</h3>

<p>While The Netherlands is just a small player, our national effort towards Open Access has made a clear
world wide impact. So has Open Science, and FAIR data, two other (but distinct) movements to improve
knowledge dissemination, and both where The Netherlands has made a significant international contribution.</p>

<p>So, Dutch universities should work harder to make the next step, after CC-BY licensing and the
<a href="https://www.openaccess.nl/nl/beleid/open-access-in-het-nederlandse-auteursrecht-taverne-regeling">Taverne</a>
green Open Access law. Let’s remove that small print of 100% Open Access that says, “but only of new
literature”. Open scientists have been working on this for long enough; it is just policy, laws, and
the mere willingness. Practically, every Dutch university can just tomorrow contact all (emeritus)
professors and make all their literature available under the Taverne rules. Many university libraries
have already started this, but this must have UNL backing with a clear mandate to make 100% of all
Dutch works available, across the full history.</p>

<p>I repeat, this is not a technical question. It is just doing it.</p>

<p>This goes for every Dutch researcher that want to have their research have the impact it can have.
Go to your university library, and send them publisher PDFs of all articles (and book chapters) you
ever published at that university. Yes, you can even email the library where you did your PhD.</p>

<p>Taverne is not the full story. Authors have to give permission, via opt-in or opt-out arrangements
with their universities. That is a problem when the researcher is no longer around. But can
we at least start with all that paywalled literature of which researchers can give that permission,
please, for the <em>full</em> appointments, going back in time? For the rest, well, we have to start
lobbying with our government. That is quite modern and should be trivial.</p>

<p>I think this is essential: you cannot fully understand the impact of an article if you do
not have access to the cited and citing literature because it is behind a paywall.</p>

<h3 id="action-2-make-the-context-of-literature-easier-to-record-and-easier-to-access">Action 2: make the context of literature easier to record and easier to access</h3>

<p>The other point Lendering made is about the self-correcting nature of science (or lack thereof).
The peer review problems were mentioned in the podcast, as well as people not easily finding literature
that disproves earlier literature. It can be a lot simpler to get informed about the status of any
research work. But right now, even something as simple as getting PubPeer and RetractionWatch
information about a References list is non-trivial. The choice of commercial publishers to favor
PDF over (scientific) HTML has been one driven by perverse incentive, not by scientific needs.
Similarly, it is trivial to learn that an article has been disproven, but not simple. University
libraries are not yet equiped correctly or funded appropriately to allow researchers (and the
general local audience who can also enter university libraries to learn stuff).</p>

<p>My gauntlet to Springer Nature about <a href="https://chem-bla-ics.linkedchemistry.info/tag/cito">CiTO citation intent annotations</a>
has yet been underwhelming, despite an in my opinion hopeful and promising pilot. I cannot complain,
and <a href="https://qlever.scholia.wiki/cito/">the interest is there</a>, but it was too difficult for Springer Nature.
And with that we are back with the one of the other main themes of the podcast.</p>

<p>With that, I isolated just a few moments from the podcast. I insist you listen to the full
podcast to get the proper context of the above quotes. If people want to learn more about
the context I see, I invite you to browse my blog (<a href="https://chem-bla-ics.linkedchemistry.info/archive/">newer posts</a>
and <a href="https://chem-bla-ics.blogspot.com/">older posts</a>).</p>

<p>Finally, thank you Krijn, thank you Jona, for this awesome 83 minutes and 16 seconds at 1.25x
speed!</p>

      <h4>References</h4>
      <ul>
      
        
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="science"/><category term="cito:citesAsRecommendedReading:10.1515/9783110782844-010"/><category term="openscience"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/08/07/scientias-oudheidkunde-wetenschap-hoe-dan.html">
      <![CDATA[ Communication of science is important to me and it comes in many formats. Educational resources, research articles, text books, popular science magazines, blogs, TV shows, podcasts, etc (in no particular order). Somewhere during the pandemic, I started listening to podcasts. I had done that, in the early days, but those early podcasts… well, let’s say, the format had not materialized yet. But also that it can be quite helpful to not play the podcast at the normal speeds, but at, say 1.25x. Podcasts, like any of the above format have a style, an audience in mind. Maybe the audience is first year students, peer in the research field, or the general audience. Or a politician that wants to get credibility. There are several podcasts that I have been listening. Not as education. The may have episode notes, but not proper scientific references. References is one thing important in the communication. Context is another. ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/scientias_episode.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/scientias_episode.png"/></entry>
  
  <entry>
    <title type="html">Molecular Inorganics: SMILES, MDL molfile v3000, and InChIs</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/07/26/inorganic-compounds-smiles-mdl-molfile-v3000-and-inchis.html" rel="alternate" type="text/html" title="Molecular Inorganics: SMILES, MDL molfile v3000, and InChIs"/>
    <published>2026-07-26T00:00:00+00:00</published>
    <updated>2026-07-26T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/yhf27-fp921</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/26/inorganic-compounds-smiles-mdl-molfile-v3000-and-inchis.html">
      <![CDATA[ <p>Making chemistry more FAIR requires unique identifiers for chemical structures. For organic compounds plenty of solutions exist that
do a great job. Last year and last week, I attended two technical <a href="https://www.inchi-trust.org/">InChI</a> meetings, both with
<a href="https://en.wikipedia.org/wiki/Organometallic_chemistry">organometallic compounds</a> and other molecular inorganics as one of the key topics. Thanks to
<a href="https://bsky.app/profile/herreslab.bsky.social">Sonja</a> (<a href="https://fed.brid.gy/bsky/herreslab.bsky.social">Mastodon bridge</a>)
and <a href="https://www.linkedin.com/in/gerd-blanke-b13115/">Gerd</a> for the invitations. My role includes thinking about what all the
work on the InChI means for the <a href="http://cdk.github.io/">Chemistry Development Kit</a>.</p>

<p>Many things came up. One was testing of new InChI functionality for these organometallic compounds, particularly the stereochemistry.
<a href="https://chem-bla-ics.linkedchemistry.info/2007/08/02/molecules-in-wikipedia.html">Wikipedia has many chemical compounds</a> and could
be a source, but <a href="https://chem-bla-ics.linkedchemistry.info/2022/11/12/wikidata-script-for-smiles-smarts-and.html">so does Wikidata</a>.
Both use the SMILES, but not all SMILES captures all the chemistry we need. And the InChI software needs
<a href="https://en.wikipedia.org/wiki/Chemical_table_file#V3000">an V3000 MDL Molfile</a>. Thanks to John and other CDK developers, there
is good support for recent cheminformatics software, but I was not sure it had what I would need.</p>

<p>This post is the first of a few related posts. This post is about converting SMILES from <a href="https://wikidata.org/">Wikidata</a>
to V3000 files. Take <a href="https://qlever.scholia.wiki/chemical/Q412415">cisplatin</a>: it has four ligands around a platinum atom,
all in a single plane:</p>

<p><img src="/assets/images/cisplatin.png" alt="" /></p>

<p>In this image, in red, is actually an annotation of how the ligands are oriented around the platinum. This is also reflected
in the <em>isomeric SMILES</em> in Wikidata: <code class="language-plaintext highlighter-rouge">Cl[Pt@SP1]([NH3])([NH3])Cl</code>.</p>

<p>The following source code is written in <a href="https://chem-bla-ics.linkedchemistry.info/tag/groovy">Groovy</a> which I have used for many
years because it is less verbose than Java. First, we set up our helper classes:</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'org.openscience.cdk'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'cdk-smiles'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'2.12'</span><span class="o">)</span>
<span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'org.openscience.cdk'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'cdk-silent'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'2.12'</span><span class="o">)</span>
<span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'org.openscience.cdk'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'cdk-ctab'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'2.12'</span><span class="o">)</span>
<span class="nd">@Grab</span><span class="o">(</span><span class="n">group</span><span class="o">=</span><span class="s1">'org.openscience.cdk'</span><span class="o">,</span> <span class="n">module</span><span class="o">=</span><span class="s1">'cdk-sdg'</span><span class="o">,</span> <span class="n">version</span><span class="o">=</span><span class="s1">'2.12'</span><span class="o">)</span>

<span class="kn">import</span> <span class="nn">org.openscience.cdk.smiles.SmilesParser</span><span class="o">;</span>
<span class="kn">import</span> <span class="nn">org.openscience.cdk.silent.SilentChemObjectBuilder</span><span class="o">;</span>
<span class="kn">import</span> <span class="nn">org.openscience.cdk.io.SDFWriter</span><span class="o">;</span>
<span class="kn">import</span> <span class="nn">org.openscience.cdk.layout.StructureDiagramGenerator</span><span class="o">;</span>
<span class="kn">import</span> <span class="nn">javax.vecmath.Vector2d</span>

<span class="n">builder</span> <span class="o">=</span> <span class="n">SilentChemObjectBuilder</span><span class="o">.</span><span class="na">getInstance</span><span class="o">()</span>
<span class="n">sp</span> <span class="o">=</span> <span class="k">new</span> <span class="n">SmilesParser</span><span class="o">(</span><span class="n">builder</span><span class="o">)</span>
<span class="n">sdg</span> <span class="o">=</span> <span class="k">new</span> <span class="n">StructureDiagramGenerator</span><span class="o">();</span>
</code></pre></div></div>

<p>With some extra code, I can actually get many compounds from Wikidata to convert to v3000 with a SAPRQL (as I have done
last year with CXSMILES and polymers, unpublished), but let’s go with a single example:</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">smiles</span> <span class="o">=</span> <span class="s2">"Cl[Pt@SP1]([NH3])([NH3])Cl"</span>
<span class="n">label</span> <span class="o">=</span> <span class="s2">"cisplatin"</span>
<span class="n">wdItem</span> <span class="o">=</span> <span class="s2">"Q412415"</span>
</code></pre></div></div>

<p>I can parse the SMILES and generated 2D coordinates with (which is also the approach by <a href="https://www.simolecule.com/cdkdepict/depict/bow/svg?smi=Cl%5BPt%40SP1%5D(%5BNH3%5D)(%5BNH3%5D)Cl&amp;zoom=2.0&amp;annotate=cip">CDK Depict</a>
which I used for the above 2D diagram of cisplatin):</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">mol</span> <span class="o">=</span> <span class="n">sp</span><span class="o">.</span><span class="na">parseSmiles</span><span class="o">(</span><span class="n">smiles</span><span class="o">)</span>
<span class="n">sdg</span><span class="o">.</span><span class="na">setMolecule</span><span class="o">(</span><span class="n">mol</span><span class="o">);</span>
<span class="n">sdg</span><span class="o">.</span><span class="na">generateCoordinates</span><span class="o">(</span><span class="k">new</span> <span class="n">Vector2d</span><span class="o">(</span><span class="mi">0</span><span class="o">,</span> <span class="mi">1</span><span class="o">));</span>
<span class="n">mol</span> <span class="o">=</span> <span class="n">sdg</span><span class="o">.</span><span class="na">getMolecule</span><span class="o">();</span>
</code></pre></div></div>

<p>If you have more than one molfile, they can be combined into a <a href="https://en.wikipedia.org/wiki/Chemical_table_file#SDF">SD file</a>,
to which additional properties can be added:</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">mol</span><span class="o">.</span><span class="na">setTitle</span><span class="o">(</span><span class="n">label</span><span class="o">)</span>
<span class="n">mol</span><span class="o">.</span><span class="na">setProperty</span><span class="o">(</span><span class="s2">"PUBCHEM_SUBSTANCE_SYNONYM"</span><span class="o">,</span> <span class="n">label</span><span class="o">)</span>
<span class="n">mol</span><span class="o">.</span><span class="na">setProperty</span><span class="o">(</span><span class="s2">"PUBCHEM_SUBSTANCE_COMMENT"</span><span class="o">,</span> <span class="n">smiles</span><span class="o">)</span>
<span class="n">mol</span><span class="o">.</span><span class="na">setProperty</span><span class="o">(</span><span class="s2">"PUBCHEM_EXT_DATASOURCE_REGID"</span><span class="o">,</span> <span class="n">wdItem</span><span class="o">)</span>
<span class="n">mol</span><span class="o">.</span><span class="na">setProperty</span><span class="o">(</span><span class="s2">"PUBCHEM_EXT_SUBSTANCE_URL"</span><span class="o">,</span> <span class="s2">"https://qlever.scholia.wiki/"</span> <span class="o">+</span> <span class="n">wdItem</span><span class="o">)</span>
</code></pre></div></div>

<p>And then generate the actual SD file with:</p>

<div class="language-groovy highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">writer</span> <span class="o">=</span> <span class="k">new</span> <span class="n">FileWriter</span><span class="o">(</span><span class="k">new</span> <span class="n">File</span><span class="o">(</span><span class="s2">"demo.sdf"</span><span class="o">))</span>
<span class="n">SDFWriter</span> <span class="n">sdfWriter</span> <span class="o">=</span> <span class="k">new</span> <span class="n">SDFWriter</span><span class="o">(</span><span class="n">writer</span><span class="o">);</span>
<span class="n">sdfWriter</span><span class="o">.</span><span class="na">getSetting</span><span class="o">(</span><span class="n">SDFWriter</span><span class="o">.</span><span class="na">OptAlwaysV3000</span><span class="o">).</span><span class="na">setSetting</span><span class="o">(</span><span class="s2">"true"</span><span class="o">);</span>
<span class="n">sdfWriter</span><span class="o">.</span><span class="na">write</span><span class="o">(</span><span class="n">mol</span><span class="o">);</span>
<span class="n">sdfWriter</span><span class="o">.</span><span class="na">close</span><span class="o">();</span>
<span class="n">writer</span><span class="o">.</span><span class="na">close</span><span class="o">();</span>
</code></pre></div></div>

<p>We then get this v3000 file:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>cisplatin
  CDK     07262618042D

  0  0  0     0  0            999 V3000
M  V30 BEGIN CTAB
M  V30 COUNTS 5 4 0 0 0
M  V30 BEGIN ATOM
M  V30 1 Cl -1.29904 2.25 0 0
M  V30 2 Pt 0 1.5 0 0
M  V30 3 N 1.29904 2.25 0 0 VAL=4
M  V30 4 N 1.29904 0.75 0 0 VAL=4
M  V30 5 Cl -1.29904 0.75 0 0
M  V30 END ATOM
M  V30 BEGIN BOND
M  V30 1 1 2 1 CFG=3
M  V30 2 1 2 3 CFG=3
M  V30 3 1 2 4 CFG=1
M  V30 4 1 2 5 CFG=1
M  V30 END BOND
M  V30 END CTAB
M  END
&gt; &lt;PUBCHEM_SUBSTANCE_COMMENT&gt;
Cl[Pt@SP1]([NH3])([NH3])Cl

&gt; &lt;PUBCHEM_EXT_DATASOURCE_REGID&gt;
Q412415

&gt; &lt;PUBCHEM_SUBSTANCE_SYNONYM&gt;
cisplatin

&gt; &lt;PUBCHEM_EXT_SUBSTANCE_URL&gt;
https://qlever.scholia.wiki/Q412415

$$$$
</code></pre></div></div>

<p>I can copy/paste the resulting v3000 content to the <a href="https://iupac-inchi.github.io/InChI-Web-Demo/">InChI Web Demo</a> to calculate the
Standard InChI:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>InChI=1S/2ClH.2H3N.Pt/h2*1H;2*1H3;/q;;;;+2/p-2
</code></pre></div></div>

<p>And this is what the two technical meetings I attended were about: <em>molecular inorganics</em>. The above InChI does not feel right,
and certainly lost the connectivity of the ligands with the platinum. However, if we add the beta option <code class="language-plaintext highlighter-rouge">-MolecularInorganics</code>,
then we get this InChI (where the <code class="language-plaintext highlighter-rouge">B</code> in <code class="language-plaintext highlighter-rouge">InChI=1B</code> reflects the beta state of this feature):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>InChI=1B/Cl2H6N2Pt/c1-5(2,3)4/h3-4H3
</code></pre></div></div>

<p>However, this beta version does not distinguish cisplatin from <a href="https://qlever.scholia.wiki/chemical/Q25403157">transplatin</a>. For that,
we need to dive into how to represent the stereochemistry of these inorganics first.</p>

      <h4>References</h4>
      <ul>
      
      
      
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="chemistry"/><category term="inchi"/><category term="wikidata"/><category term="smiles"/><category term="pubchem"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/26/inorganic-compounds-smiles-mdl-molfile-v3000-and-inchis.html">
      <![CDATA[ Making chemistry more FAIR requires unique identifiers for chemical structures. For organic compounds plenty of solutions exist that do a great job. Last year and last week, I attended two technical InChI meetings, both with organometallic compounds and other molecular inorganics as one of the key topics. Thanks to Sonja (Mastodon bridge) and Gerd for the invitations. My role includes thinking about what all the work on the InChI means for the Chemistry Development Kit. ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cisplatin.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/cisplatin.png"/></entry>
  
  <entry>
    <title type="html">Carbon beats gold: Diamond Open Access is the future #2</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future2.html" rel="alternate" type="text/html" title="Carbon beats gold: Diamond Open Access is the future #2"/>
    <published>2026-07-13T00:00:00+00:00</published>
    <updated>2026-07-13T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/gfrxs-bxn43</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future2.html">
      <![CDATA[ <p>We need a lot more than <a href="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future.html">diamond open access</a>
to really improve the publishing models. That said, but there are <a href="https://chem-bla-ics.linkedchemistry.info/2024/09/16/publishing.html">examples that</a>
<a href="https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html">diamond open access publishers</a> actually want to
improve more just the access to the knowledge dissemination infrastructure.
But infrastructure is not only technologies; it also includes the many social aspects that are involved in adoption.
And we saw <a href="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future.html">enough of that in the open access transition</a>.</p>

<p>One aspect that limits the uptake of diamond open access models is that the scholarly community needs to change
their trust model. We have seen a transition from societal publishers to commercial publishers, from scholarly-led
to publisher-led. And commercial publishers certainly <a href="https://chem-bla-ics.linkedchemistry.info/2021/06/11/conflict-of-interest-or-why-i-am.html">breached my trust</a>
(see also doi:<a href="http://doi.org/10.5281/ZENODO.4926030">10.5281/zenodo.4926030</a>).</p>

<p>Many scholars prefer the certainty of a scholarly journal where they can trust that the editorial board and
their reviewers take their job seriously. This matters. The notion that peer reviewers may reject your work is an
extrinsic motivation for at least some researchers to do a better job (personal communication). And that feature
works well: there are indications that research published as preprint sees limited change after the formal
peer review process (doi:<a href="https://doi.org/10.1038/d41586-026-02167-3">10.1038/d41586-026-02167-3</a> and
doi:<a href="https://doi.org/10.64898/2026.06.30.735556">10.64898/2026.06.30.735556</a>):</p>

<blockquote>
  <p>“For most scientists, reputation is important, and they would not upload something they would not be comfortable
with publishing in a journal later.”</p>
</blockquote>

<p>In an ideal situation, reputation would not be part of the whole equation: researchers never submit anything they think
is not scientifically sound. Journals would never desk-reject an article that is scientifically sound. The question
is mostly how to reach that. Everyone is busy, and it has not been so long ago that the journal impact factor was
by many as indicator of scientific quality. The reality is far more complex than that.</p>

<h2 id="finding-diamond-open-access-journals">Finding Diamond Open Access journals</h2>

<p>So, let’s assume you are interested in the idea, and want to know what diamond open access journals exist in your
research field, before you can even consider the quality of that journal, you need to know which journal is diamond
open access. Therefore, I <a href="https://mastodon.social/@egonw/116906626546588093">asked yesterday on Mastodon</a> about databases that list
which journal are diamond.</p>

<p>It actually turns out that the <a href="https://doaj.org/">Directory of Open Access Journals</a> (DOAJ) does not indicate
if a journal is diamond. You can filter on zero APC, but that is not the same. In fact, zero APC is routinely used
by big publishers to launch new journals. One they have critical mass, they will start charging non-zero APCs.</p>

<p>I already received some replies about possible options:</p>

<ol>
  <li><a href="https://ddh.edch.eu/en">Diamond Discovery Hub</a>: lists just over 4,000 diamond journals, but seems closed data (thx <a href="https://bsky.app/profile/najko.bsky.social/post/3mqh25otvvk2z">Najko</a>! see also doi:<a href="https://doi.org/10.15291/libellarium.4569">10.15291/libellarium.4569</a>)</li>
  <li><a href="https://wheretopublish.github.io/">wheretopublish.github.io</a>: open data, accepts additions, but lists only 9 diamond journals (thx <a href="https://bsky.app/profile/annecmg.bsky.social/post/3mqj4envqns24">Anne</a>!)</li>
  <li><a href="https://service.tib.eu/bison/">B!SON</a>: a recommender that depends on DOAJ, so also without clear diamond indication (thx <a href="https://degrowth.social/@yala/116906904947384383">Jon</a>!)</li>
</ol>

<p>All pieces of the puzzle.</p>

<p>When I grew up as a scholar, I learned to publish in journals where you respect the research, where you see important
research in your field get published. Nowadays, this is changing. In fact, young researchers (tho many older scholars alike)
fear to publish in predatory journal by accident (really! <code class="language-plaintext highlighter-rouge">[citation_pending]</code>). The big publishers and the new
publishers have been adding so many new journals, and academia has really pushed the “publish as many articles
as you can” model, that it kind of makes sense why things are different now.</p>

<p>And, a central database with trustworthy information about the quality of diamond open access journals
is just needed.</p>

<p>A final note, while I have not been able to create a list of diamond open access journals with OpenAlex,
already in February 2025 I found a way to list <a href="https://edu.nl/q3uf3">articles tagged with a <em>diamond status</em></a>.
This query lists over three thousand of such annoated articles for Maastricht University:</p>

<p><img src="/assets/images/openalex_diamondOA_UM_articles.png" alt="" /></p>

<p>(Interestingly, the journal where the three most cited articles were published turned diamond already in 2004
and was discontinued about half a year go “Due to recent changes in operational resources”,
<a href="https://en.wikipedia.org/wiki/Environmental_Health_Perspectives">according to Wikipedia</a>.)</p>

<h2 id="building-trust">Building trust</h2>

<p>So, while the above solution do allow you to increase your success rate or finding a diamond open access journal
in your field, the next question is: what can you learn about this journal. For example, what does Wikipedia
say about the journal (if anything)? Do you know people that have published in that journal? The OpenAlex
screenshot shows</p>

<p>That make me wonder what <a href="https://www.wikidata.org/">Wikidata</a> has to offer. After all, as one of the largest
linked open data sources, they must have some information.</p>

<p>I am not disappointed. The <em>type</em> <a href="https://www.wikidata.org/wiki/Q108440863">diamond open-access journal</a> is actively
used, and there is plenty of other of useful fields for these journals. I came up with
<a href="https://edu.nl/xh7h9">this SPARQL query</a> (with some QLever sugar for labels):</p>

<div class="language-sparql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;http://www.wikidata.org/prop/direct/&gt;</span><span class="w">
</span><span class="k">PREFIX</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;http://www.wikidata.org/entity/&gt;</span><span class="w">
</span><span class="k">PREFIX</span><span class="w"> </span><span class="nn">rdfs</span><span class="o">:</span><span class="w"> </span><span class="nn">&lt;http://www.w3.org/2000/01/rdf-schema#&gt;</span><span class="w">
</span><span class="k">SELECT</span><span class="w"> </span><span class="k">DISTINCT</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nv">?diamondjournalLabel</span><span class="w">
      </span><span class="c1"># (SAMPLE(?url_) AS ?website)</span><span class="w">
      </span><span class="p">(</span><span class="nb">SAMPLE</span><span class="p">(</span><span class="nv">?issnl_</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?issnl</span><span class="p">)</span><span class="w">
      </span><span class="p">(</span><span class="nb">GROUP_CONCAT</span><span class="p">(</span><span class="k">DISTINCT</span><span class="w"> </span><span class="nb">STR</span><span class="p">(</span><span class="nv">?issn_</span><span class="p">))</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?issn</span><span class="p">)</span><span class="w">
      </span><span class="p">(</span><span class="nb">SAMPLE</span><span class="p">(</span><span class="nv">?lang_</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?language</span><span class="p">)</span><span class="w">
      </span><span class="p">(</span><span class="nb">SAMPLE</span><span class="p">(</span><span class="nv">?doaj_</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?doaj</span><span class="p">)</span><span class="w">
      </span><span class="p">(</span><span class="nb">GROUP_CONCAT</span><span class="p">(</span><span class="k">DISTINCT</span><span class="w"> </span><span class="nb">STR</span><span class="p">(</span><span class="nv">?subject</span><span class="p">);</span><span class="nb">separator</span><span class="p">=</span><span class="s2">", "</span><span class="p">)</span><span class="w"> </span><span class="k">AS</span><span class="w"> </span><span class="nv">?subjects</span><span class="p">)</span><span class="w">
</span><span class="k">WHERE</span><span class="w"> </span><span class="p">{</span><span class="w">
  </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P31</span><span class="w"> </span><span class="nn">wd</span><span class="o">:</span><span class="ss">Q108440863</span><span class="w"> </span><span class="p">.</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P236</span><span class="w"> </span><span class="nv">?issn_</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P7363</span><span class="w"> </span><span class="nv">?issnl_</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P856</span><span class="w"> </span><span class="nv">?url_</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P407</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="err">@en@</span><span class="nn">rdfs</span><span class="o">:</span><span class="ss">label</span><span class="w"> </span><span class="nv">?lang_</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P5115</span><span class="w"> </span><span class="nv">?doaj_</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="k">OPTIONAL</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nn">wdt</span><span class="o">:</span><span class="ss">P921</span><span class="w"> </span><span class="o">/</span><span class="w"> </span><span class="err">@en@</span><span class="nn">rdfs</span><span class="o">:</span><span class="ss">label</span><span class="w"> </span><span class="nv">?subject</span><span class="w"> </span><span class="p">}</span><span class="w">
  </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="err">@en@</span><span class="nn">rdfs</span><span class="o">:</span><span class="ss">label</span><span class="w"> </span><span class="nv">?diamondjournalLabel</span><span class="w"> </span><span class="p">.</span><span class="w">
</span><span class="p">}</span><span class="w"> </span><span class="k">GROUP</span><span class="w"> </span><span class="k">BY</span><span class="w"> </span><span class="nv">?diamondJournal</span><span class="w"> </span><span class="nv">?diamondjournalLabel</span><span class="w">
</span></code></pre></div></div>

<p>This lists more than 1,500 diamond open access journals:</p>

<p><img src="/assets/images/wikidata_diamondOA_journals.png" alt="" /></p>

<p>With Wikidata we can then <a href="https://edu.nl/ekbq9">list UM authors that published in diamond journals</a>.
That list is remarkably short, however, if you compare it with the earlier results from OpenAlex. That is because
Wikidata only contains a subset of scientific literature and not all articles are correctly linked to the items
about their authors. But I also spotted journals in the OpenAlex list that are not really diamond.</p>

<h2 id="what-is-next">What is next?</h2>

<p>Clearly, we need better metadata, but things are not looking that bad either! We have something to work with.
What is your favorite diamond open access journal?</p>

      <h4>References</h4>
      <ul>
      
        
      
        
      
        
      
        
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="openaccess"/><category term="mycito:citesAsRecommendedReading:10.5281/ZENODO.4926030"/><category term="cito:includesQuotationFrom:10.1038/d41586-026-02167-3"/><category term="cito:citesAsDatasource:10.64898/2026.06.30.735556"/><category term="cito:citesForInformation:10.15291/libellarium.4569"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future2.html">
      <![CDATA[ We need a lot more than diamond open access to really improve the publishing models. That said, but there are examples that diamond open access publishers actually want to improve more just the access to the knowledge dissemination infrastructure. But infrastructure is not only technologies; it also includes the many social aspects that are involved in adoption. And we saw enough of that in the open access transition. ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/wikidata_diamondOA_journals.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/wikidata_diamondOA_journals.png"/></entry>
  
  <entry>
    <title type="html">Carbon beats gold: Diamond Open Access is the future #1</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future.html" rel="alternate" type="text/html" title="Carbon beats gold: Diamond Open Access is the future #1"/>
    <published>2026-07-13T00:00:00+00:00</published>
    <updated>2026-07-13T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/bsnp5-68v45</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future.html">
      <![CDATA[ <p>Thirty years ago, researchers were struggling getting access to literature they wanted to read. I remember PhD candidates
visiting friends at nearby universities for a meetup, and while there, for the copying machine in the remote library. Faster
and cheaper than inter-library loaning. Scholarly journals were still printed on paper and distributed to university
libraries. That was expensive. Therefore, many libraries provided only access to a subset of journals.</p>

<p>Then the internet came. Big publishers started sharing offering digital-only subscriptions. This would lower the cost
and therefor the prize. That was the idea. But the big publishers remained expensive publishers.</p>

<p>Then open access came. The idea is that journal article would have a open license, often CC-BY or similar. This meant
that researchers could share articles with colleagues and students without additional cost. The idea was that this
would lower the cost of the journals. After all, libraries could reshare the articles, and contribute in funding
the distribution of the knowledge. But the big publishers insisted to be the only source of the PDFs, and they
remained expensive publishers.</p>

<p>Then came the article-processing charge (APC) model. The authors would contribute for the publishing and distribution
process. For some years, new publishers showed, organized by and for scholars, keeping the APC close to the cost
of production and distribution. The APC was estimated at somewhere between 50 and 600 euro per article, likely more
now, after the inflation of the last five years. But scholars did not trust these publishers. Second, publishers
saw a market and a business model in APC-based publishing. Libraries were pressured to keep access to big publishers
that realized that if scholars wanted to keep publishing there, they could charge higher APC based on popularity.
New expensive publishers entered the market, some with good intentions, some with bad intentions. And among all the
confusion, the big publishers remained expensive publishers.</p>

<p>This is where we are now. Because sharing knowledge was a great idea, we want to keep the license, but we need
to drop the financial incentives: we need to drop the APC.</p>

<p>Enter the world of <a href="https://en.wikipedia.org/wiki/Diamond_open_access">diamond open access</a>, where science is
published under an open license, and neither the reviewers, not the readers, nor the authors pay, all key stakeholder
in the exchange of scientific knowledge. And no tax by publishers.</p>

<p>Can we do this? More and more people think so, and diamond open access is gaining traction.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="openaccess"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/07/13/carbon-beats-gold-diamond-open-access-is-the-future.html">
      <![CDATA[ Thirty years ago, researchers were struggling getting access to literature they wanted to read. I remember PhD candidates visiting friends at nearby universities for a meetup, and while there, for the copying machine in the remote library. Faster and cheaper than inter-library loaning. Scholarly journals were still printed on paper and distributed to university libraries. That was expensive. Therefore, many libraries provided only access to a subset of journals. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Schema-Driven Interoperability of Biomedical Data Resources for Improved Knowledge Dissemination and Reuse (SchemaInterop)</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/06/15/schema-driven-interoperability.html" rel="alternate" type="text/html" title="Schema-Driven Interoperability of Biomedical Data Resources for Improved Knowledge Dissemination and Reuse (SchemaInterop)"/>
    <published>2026-06-15T00:00:00+00:00</published>
    <updated>2026-06-15T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/yd794-47y39</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/06/15/schema-driven-interoperability.html">
      <![CDATA[ <p>Today starts a new project. NWO’s <a href="https://www.openscience.nl/">Open Science NL</a> awarded <a href="http://orcid.org/0000-0002-4904-3269">Dr Tooba Abbassi-Daloii</a>
(Amsterdam UMC) with a grant to work on the interoperability between biomedical databases which all have different scope and design
(doi:<a href="https://doi.org/10.61686/AQNSM35060">10.61686/AQNSM35060</a>)). This project will create an open, interoperable framework to connect these databases,
enabling efficient access and data use. By applying schema harmonization, aligning database designs, and using open science approaches, it promotes data
transparency, reusability, and interoperability. Our approach implements the FAIR principles, ensuring analysis across databases is accurate and
transparent, e.g., when combining gene expression data with gene variant knowledge. The project brings together researchers from the Amsterdam UMC,
Universiteit Maastricht, and Leiden University Medical Center.</p>

<p>The full proposal can be read in the Open Science NL Collection on Zenodo (doi:<a href="https://doi.org/10.5281/zenodo.19704231">10.5281/zenodo.19704231</a>).</p>

      <h4>References</h4>
      <ul>
      
        <li><a href="https://doi.org/10.5281/ZENODO.19704231">10.5281/ZENODO.19704231</a></li>
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="openscience"/><category term="doi:10.5281/ZENODO.19704231"/><category term="schemainterop"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/06/15/schema-driven-interoperability.html">
      <![CDATA[ Today starts a new project. NWO’s Open Science NL awarded Dr Tooba Abbassi-Daloii (Amsterdam UMC) with a grant to work on the interoperability between biomedical databases which all have different scope and design (doi:10.61686/AQNSM35060)). This project will create an open, interoperable framework to connect these databases, enabling efficient access and data use. By applying schema harmonization, aligning database designs, and using open science approaches, it promotes data transparency, reusability, and interoperability. Our approach implements the FAIR principles, ensuring analysis across databases is accurate and transparent, e.g., when combining gene expression data with gene variant knowledge. The project brings together researchers from the Amsterdam UMC, Universiteit Maastricht, and Leiden University Medical Center. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">The launch of the Virtual Human Platform</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/06/06/the-launch-of-the-virtual-human-platform.html" rel="alternate" type="text/html" title="The launch of the Virtual Human Platform"/>
    <published>2026-06-06T00:00:00+00:00</published>
    <updated>2026-06-06T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/s784p-s1y68</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/06/06/the-launch-of-the-virtual-human-platform.html">
      <![CDATA[ <p><img src="/assets/images/vhp_platform.png" style="width: 30%; display: block; margin-left: auto; margin-right: auto; float: right" alt="Screenshot of the Virtual Human Platform website, showing a logo, three section panels (Case Studies, Tools, Methods, Data), and a short description. The page is just the top part and includes several menus at the top." />
Nine days ago, the <a href="https://vhp4safety.nl/">VHP4Safety</a> project
(see <a href="https://chem-bla-ics.linkedchemistry.info/tag/vhp4safety">these posts</a>)
held a launch event in Utrecht for the
<a href="https://platform.vhp4safety.nl/">Virtual Human Platform</a> (VHP), a key result of the
<a href="https://www.nwo.nl/en/researchprogrammes/dutch-research-agenda-nwa">Dutch Research Agenda</a> (NWA,
from the Dutch <em>Nationale Wetenschapsagenda</em>). Despite the name, the NWA is just one part
of the NWO funding mechanisms, but like the <a href="https://www.openscience.nl/en/about-us">NWO Open Science programme</a>
it is funding with a specific purpose. And the purpose of the NWA is to answer
research and societal questions that the Dutch people together defined and a public
consultation (many years ago). VHP4Safety is answering to one of those questions.</p>

<h2 id="co-creation">Co-creation</h2>

<p>The project is still running another few months, but the <a href="https://www.sciencrew.com/c/9347/a/335652577?title=Launch_of_the_Virtual_Human_Platform">launch last week</a>
gives us the opportunity to include feedback from the stakeholders from the Dutch
society, many of which have been involved in the project via designathons and
hackathons (see doi:<a href="https://doi.org/10.14573/altex.2407211">10.14573/altex.2407211</a>).</p>

<p>The VHP4Safety platform is a co-creation created by most of the people working
on the VHP4Safety grant. Some people focused on innovation and education (RL3),
others on the regulatory questions (RL2), and some on the development of the
platform (RL1). The research line 1 (RL1) included a work package on the
technological development, work package 1.1, and that was led by Maastricht
University (Ozan and me) and the Applied University of Utrecht (Dr. Marc Teunis).
This project would not be together without the leadership by
Prof. dr. ir. Juliette Legler, Dr. Cyrille Krul, and Prof. dr. Anne Kienhuis
 (see <a href="https://video.edu.nl/w/rvUKc7J4E4HEt2TEbJEQt9">this video</a>).</p>

<h2 id="not-just-technology">Not just technology</h2>

<p>I have to give a huge shout out to Ozan whom had the daunting task
to set up something like OpenRiskNet (doi:<a href="https://doi.org/10.1016/j.toxlet.2018.06.617">10.1016/j.toxlet.2018.06.617</a>),
a project with at least twice as much
funding for operating and documenting just the technical platform, but also
help other partners getting their work on the platform. Also shout outs to 
Luc who in our group first explored how to translate the OpenRiskNet platform
to VHP4Safety with <a href="https://en.wikipedia.org/wiki/Kubernetes">kubernetes</a> from
which we concluded that that was not an option for us. And to Sean in our group
who introduced us to <a href="https://www.geeksforgeeks.org/devops/introduction-to-docker-swarm-mode/">Docker Swarm</a>.</p>

<p>But that is just one aspect of the technological layers. The design outlined
in the original proposal is based on earlier projects, including OpenRiskNet,
eNanoMapper, OpenRiskNet, Open PHACTS, NanoCommons, SbD4Nano and many others.
It is based on open standards developed and/or adopted by
<a href="https://elixir-europe.org/">ELIXIR Europe</a> projects and many other organisations.</p>

<p>And then we have not even covered the content on the platform.</p>

<h2 id="a-virtual-human">A virtual human</h2>

<p>Building full virtual human is an ambition. Many <a href="https://en.wikipedia.org/wiki/Digital_twin">digital twins</a>
capture just one part of human biology. For safety assessment we need many models,
data from experiments, and knowledge bases. And we need a clear narrative that
describes how those isolated solutions are integrated so that regulatory questions
can be answered. That co-created combination is the launched <em>virtual human platform</em>.</p>

<p>Underlying the VHP is a good bit of open science, though it also integrated 
proprietary solutions, currently needed to be able to replace animal testing.
And our modular co-creation resulted in <a href="https://github.com/VHP4Safety">many separate git repositories</a>.
This allows distributed development models and all contributors to take ownership
of the development of their contributions. Marc and Frank introduced a agile computing
approach that we adopted to guide the development of the full platform.</p>

<p>There is so much to write up about the platform (and we will), but for now I want
to highlight a few essential git repositories underlying the 1.0 version of the
platform we launched last week (along with the number of contributors in the past two years):</p>

<ul>
  <li><a href="https://github.com/VHP4Safety/virtual-human-platform">virtual-human-platform</a> (<a href="https://github.com/VHP4Safety/virtual-human-platform/graphs/contributors?from=6%2F1%2F2024">10 contributors</a>): software that provide the platform UX</li>
  <li><a href="https://github.com/VHP4Safety/cloud">cloud</a> (<a href="https://github.com/VHP4Safety/cloud/graphs/contributors?from=6%2F1%2F2024">11 contributors</a>): collects the meta data about (computational) services</li>
  <li><a href="https://github.com/VHP4Safety/ui-casestudy-config">ui-casestudy-config</a> (<a href="https://github.com/VHP4Safety/ui-casestudy-config/graphs/contributors?from=5%2F31%2F2025">7 contributors</a>): collects the details of the narratives of the case studies</li>
</ul>

<p>This includes <a href="https://github.com/aniekdewinter">Aniek</a>, <a href="https://github.com/FW94">Fabian</a>,
<a href="https://github.com/iaortega">Isaac</a>, <a href="https://github.com/senseibelbi">Ivo</a>,
<a href="https://github.com/jmillanacosta">Javier</a>, <a href="https://github.com/johannehouweling">Jente</a>,
<a href="https://github.com/LindeSchoenmaker">Linde</a>, <a href="https://github.com/marvinm2">Marvin</a>,
<a href="https://github.com/mirthhe">Myrthe</a>, <a href="https://github.com/saadlodhi0916">Saad</a>,
<a href="https://github.com/ShakiraPortfolio">Shakira</a>, and <a href="https://github.com/youphendriks">Youp</a>,
in addition to the earlier named <a href="https://github.com/ozancinar">Ozan</a> and
<a href="https://github.com/Maddocent">Marc</a>.</p>

<p>This excludes the many contributions via the designathons and hackathons that are
behind many of the commits to these repositories. And this also excludes the
many <a href="https://github.com/VHP4Safety/">other source code repositories</a> for tools
and services developed and made available on this VHP, with even more researchers.</p>

<p>Very much aware of all the work that is still ahead of us, I am happy with the
important milestone of this release. Thank you to
<a href="https://github.com/orgs/VHP4Safety/people">everyone who contributed to this co-creation</a>,
all the <a href="https://www.sciencrew.com/c/9319/a/329042452?title=VHP4Safety_Partners">involved institutes</a>,
NWO for <a href="https://www.nwo.nl/en/projects/nwa129219272">the funding</a>, and the Dutch
public for the important NWA question.</p>

<p>The project delivered!</p>

      <h4>References</h4>
      <ul>
      
      
      
      
        <li><a href="https://doi.org/10.14573/ALTEX.2407211">10.14573/ALTEX.2407211</a></li>
      
        <li><a href="https://doi.org/10.1016/j.toxlet.2018.06.617">10.1016/j.toxlet.2018.06.617</a></li>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="vhp4safety"/><category term="bioschemas"/><category term="openscience"/><category term="elixir"/><category term="doi:10.14573/ALTEX.2407211"/><category term="justdoi:10.1016/j.toxlet.2018.06.617"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/06/06/the-launch-of-the-virtual-human-platform.html">
      <![CDATA[ Nine days ago, the VHP4Safety project (see these posts) held a launch event in Utrecht for the Virtual Human Platform (VHP), a key result of the Dutch Research Agenda (NWA, from the Dutch Nationale Wetenschapsagenda). Despite the name, the NWA is just one part of the NWO funding mechanisms, but like the NWO Open Science programme it is funding with a specific purpose. And the purpose of the NWA is to answer research and societal questions that the Dutch people together defined and a public consultation (many years ago). VHP4Safety is answering to one of those questions. ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/vhp_platform.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/vhp_platform.png"/></entry>
  
  <entry>
    <title type="html">New paper: pyBiodatafuse: Extending interoperability of data using modular queries across biomedical resources</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/05/30/new-paper-pybiodatafuse.html" rel="alternate" type="text/html" title="New paper: pyBiodatafuse: Extending interoperability of data using modular queries across biomedical resources"/>
    <published>2026-05-30T00:00:00+00:00</published>
    <updated>2026-05-30T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/7n2bs-zsm80</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/30/new-paper-pybiodatafuse.html">
      <![CDATA[ <p>The number of data and knowledge source relevant to your biological or chemical question
increases every year. They all come with different API and different data models. These
need to be documented and mapped. What better way to do that than actually do that and
then use that. I never asked, but I can imagine that was the original idea of Tooba
and Yojana. At the very least, it demonstrates the level of interoperability we need
in the life sciences.</p>

<p>In a recent paper, <a href="https://orcid.org/0000-0002-7683-0452">Yojana Gadiya</a>,
<a href="https://orcid.org/0000-0002-4166-7093">Javier Millán Acosta</a>, and
<a href="https://orcid.org/0000-0002-4904-3269">Tooba Abbassi-Daloii</a> led a project called
BioDataFuse (worked on at the biohackathons of ELIXIR in <a href="https://doi.org/10.37044/osf.io/mhsqp">2023</a>
and <a href="https://doi.org/10.37044/osf.io/ptmg5_v1">2024</a>
and of SWAT4HCLS in <a href="https://ceur-ws.org/Vol-3890/paper-23.pdf">2024</a>
and <a href="https://ceur-ws.org/Vol-4196/paper_71.pdf">2025</a>) and the matching Python package,
<a href="https://github.com/BioDataFuse/pyBiodatafuse">pyBiodatafuse</a>
(doi:<a href="https://doi.org/10.1093/bioinformatics/btag064">10.1093/bioinformatics/btag064</a>).</p>

<p>With a group of researchers from The Netherlands, Switzerland, Czech Republic, and
the USA, multiple databases are wrapped in a uniform data model. The package
allows the generation of a graph across the imported databases which can then
be further analyzed and visualized. This is an example (RDF) graph that was generated:</p>

<p><img src="/assets/images/pyBiodatafuseGraph.png" alt="" /></p>

<p>Seeing this kind of interoperability brings back <a href="https://chem-bla-ics.linkedchemistry.info/2010/03/04/rdf-jena-bioclipse-eclipse-zest-2-icons.html">good memories</a>.</p>

<p>Congrats to all authors!</p>

      <h4>References</h4>
      <ul>
      
      
        <li><a href="https://doi.org/10.1093/BIOINFORMATICS/BTAG064">10.1093/BIOINFORMATICS/BTAG064</a></li>
      
        <li><a href="https://doi.org/10.37044/OSF.IO/MHSQP">10.37044/OSF.IO/MHSQP</a></li>
      
        <li><a href="https://doi.org/10.37044/OSF.IO/PTMG5_V1">10.37044/OSF.IO/PTMG5_V1</a></li>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="python"/><category term="data"/><category term="doi:10.1093/BIOINFORMATICS/BTAG064"/><category term="doi:10.37044/OSF.IO/MHSQP"/><category term="justdoi:10.37044/OSF.IO/PTMG5_V1"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/30/new-paper-pybiodatafuse.html">
      <![CDATA[ The number of data and knowledge source relevant to your biological or chemical question increases every year. They all come with different API and different data models. These need to be documented and mapped. What better way to do that than actually do that and then use that. I never asked, but I can imagine that was the original idea of Tooba and Yojana. At the very least, it demonstrates the level of interoperability we need in the life sciences. ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pyBiodatafuseGraph.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/pyBiodatafuseGraph.png"/></entry>
  
  <entry>
    <title type="html">FAIR Implementation Profiles: Chemistry</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/05/05/fair-implementation-profiles.html" rel="alternate" type="text/html" title="FAIR Implementation Profiles: Chemistry"/>
    <published>2026-05-05T00:00:00+00:00</published>
    <updated>2026-05-05T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/x115q-m5f95</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/05/fair-implementation-profiles.html">
      <![CDATA[ <p>I have had this on my todo list for way too long: writing about <a href="https://www.go-fair.org/how-to-go-fair/fair-implementation-profile/">FAIR Implementation Profiles</a>,
or FIPs for short (see also doi:<a href="https://doi.org/10.1007/978-3-030-65847-2_13">10.1007/978-3-030-65847-2_13</a>):</p>

<blockquote>
  <p>The FIP is a collection of FAIR implementation choices made by a community of practice for each of the FAIR Principles.</p>
</blockquote>

<p>In the early GO FAIR days, people referred to the <em>challenges and choices</em>. FIPs are a formal approach to report the choices.
The days of the <em>implementation networks</em>, like the Chemistry Implementation Network (doi:<a href="https://doi.org/10.1162/dint_a_00035">10.1162/dint_a_00035</a>),
or the AdvancedNano network (doi:<a href="https://doi.org/10.1016/j.impact.2024.100513">10.1016/j.impact.2024.100513</a>). For the first,
we absolutely agreed on the InChI, for example, but we never wrote that down as a FIP. But even without FIPs, the
idea was that communities could learn from each other, could converge. This is where the term
<a href="https://www.go-fair.org/today/fair-matrix/">FAIR Convergence Matrix</a> comes from. (I had forgotten I was actually
part of the matching Working Group. If only I had dedicated funding for it at the time.)</p>

<p>Now, I had on my wishlist to write up a list over FIPs around chemistry. This is relevant for the
<a href="https://video.edu.nl/w/pjf5vFFU287AGfYA2SXpE5">FAIR4ChemNL</a> project, but obviously beyond that. And while we did
a lot of FAIRification work in, particularly, the nanosafety cluster projects (eNanoMapper, NanoCommons, NanoSolveIT,
RiskGONE, and SbD4Nano), a lot of this never was formalized as FIPs. Second, various relevant standards have been
proposed in chemistry, for <a href="https://doi.org/10.1186/s13321-021-00520-4">chemical compounds</a> and
<a href="https://doi.org/10.1021/acsenvironau.5c00314">transformation products</a>, among plenty of other things.
But during the first <a href="https://elixir-europe.org/communities/toxicology">ELIXIR Toxicology Community</a>
workshop in Utrecht, we had FIPs on the agenda too (see <a href="https://doi.org/10.37044/osf.io/un2rw">this report</a>).</p>

<h2 id="fips-in-chemistry">FIPs in Chemistry</h2>

<p>But while I had already several browser tabs open for months, I never could find the courage to start making the list. Silly.
Therefore, this list should be considered a start. I don’t think it is exhaustive. I know there is an index somewhere,
but I cannot find back that browser tab. Additions are most welcome: I will update this post (thanks to git and Rogue Scholar).
In brackets I list (a selection of) the FAIR enabling resources mentioned in that FIP.</p>

<ul>
  <li><a href="https://fip-wizard.ds-wizard.org/wizard/projects/2f1c0e80-3c9d-4967-9dc6-fd05fac96269/metrics">Adverse Outcome Pathways FIP</a> (DOI, AOP Wiki ID)</li>
  <li><a href="https://docs.google.com/spreadsheets/d/1yNEYzJRbx10RkuqJmLcO7RPOq-nrimUKQLQW6uWzfho/edit?gid=127295437#gid=127295437">Toxicogenomics</a> (ISA-Tab, BioStudies, ArrayExpress)</li>
  <li><a href="https://fip-wizard.ds-wizard.org/wizard/projects/b0d84171-4556-46ed-9ace-c1c58db38092">WorldFAIR WP04 NANOMATERIALS FIP01</a> (DOI, ORCID, UUID, ROR, InChI, DataCite, QMRF, RDF/JSON-LD, OWL; doi<a href="https://doi.org/10.5281/zenodo.7378109">10.5281/zenodo.7378109</a>)</li>
</ul>

<p>Is that all? No, I do not think so, but this is all I can easily find (ironically). Well, I guess you now see why I have had a so much trouble starting to write down this post…
Someone will surely give me some pointers…</p>

<h2 id="ackknowledgments">Ackknowledgments</h2>

<p>I also like to thank all the people involved in these FIPs. I strongly encourage you to look up all the people who
wrote them down. I also like to thank the INTOXICOM workshop FIP experts Iseult Lynch and Gerhard Burger for
sharing their knowledge.</p>

<p>Finally, I recommend checking out the <a href="https://fip.fair-wizard.com/">FIP Wizard</a>.</p>

      <h4>References</h4>
      <ul>
      
        <li><a href="https://doi.org/10.1162/dint_a_00035">10.1162/dint_a_00035</a></li>
      
        <li><a href="https://doi.org/10.1016/J.IMPACT.2024.100513">10.1016/J.IMPACT.2024.100513</a></li>
      
        
      
        
      
        <li><a href="https://doi.org/10.37044/OSF.IO/UN2RW">10.37044/OSF.IO/UN2RW</a></li>
      
        
      
        <li><a href="https://doi.org/10.5281/zenodo.7378109">10.5281/zenodo.7378109</a></li>
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="fair"/><category term="doi:10.1162/dint_a_00035"/><category term="doi:10.1016/J.IMPACT.2024.100513"/><category term="cito:citesAsPotentialSolution:10.1186/s13321-021-00520-4"/><category term="cito:citesAsPotentialSolution:10.1021/acsenvironau.5c00314"/><category term="doi:10.37044/OSF.IO/UN2RW"/><category term="cito:citesForInformation:10.1007/978-3-030-65847-2_13"/><category term="justdoi:10.5281/zenodo.7378109"/><category term="fair4chemnl"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/05/fair-implementation-profiles.html">
      <![CDATA[ I have had this on my todo list for way too long: writing about FAIR Implementation Profiles, or FIPs for short (see also doi:10.1007/978-3-030-65847-2_13): ]]>
    </summary></entry>
  
  <entry>
    <title type="html">One Million IUPAC names #5: a new approach and 400k names</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/05/02/one-million-iupac-names-5-a-new-approach.html" rel="alternate" type="text/html" title="One Million IUPAC names #5: a new approach and 400k names"/>
    <published>2026-05-02T00:00:00+00:00</published>
    <updated>2026-05-02T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/gqtbx-jta57</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/02/one-million-iupac-names-5-a-new-approach.html">
      <![CDATA[ <p>About fifteen months ago a new project started: <a href="https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html">One Million IUPAC names</a>:</p>

<blockquote>
  <p>Thus, the idea came up, can we create a set of 1 million unique IUPAC names found in literature?</p>
</blockquote>

<p>We started out with using <a href="https://europepmc.org/">Europe PMC</a> to get JATS XML files for the full texts of open access articles.
Parsing the XML is easy and the text paragraphs are passed through OSCAR and OPSIN. That has not changed.</p>

<p>What did change last weekend is something I had long on my todo list (but life interfered). The first approach
was to ask for named entities using the Europe PMC APIs. But I quickly realized that with OSCAR and OPSIN we could
get more names out of the articles. The next step was to move from Google Colab to a command line script.
That gave another boost, as explained in <a href="http://localhost:4000/2025/04/27/one-million-iupac-names-2-the-100-thousand-milestone.html">this second post in the series</a>.
We reached 200 thousand names in <a href="https://chem-bla-ics.linkedchemistry.info/2025/06/09/one-million-iupac-names.html">june 2025</a>
but then things slowed down again in the growth. <a href="https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.html">Two months later</a>
we only had 75 thousand more. However, plenty of discussion was happening and there turned out to be
other, larger collections of IUPAC names under an open license. Millions of names, actually.</p>

<p>But another problem emerged. We were still using the Europe PMC API and were basically asking for open access
articles between two dates. Practically, the API could answer requests between 1 and max 3 days. Beyond that,
times outs and 404s became an issue. Moreover, because these dates are publications dates and not the dates
on which the JATS were deposited, I had to got back to previous months and redo the queries. That gave another
5 thousand names since last August. Something had to change.</p>

<h2 id="the-new-approach">The new Approach</h2>

<p>Europe PMC, however, also provides the JATS XML files as download on <a href="https://europepmc.org/ftp/oa/">their FTP site</a>.
Already that <a href="https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.html">august 2025</a> I had
a prototype and knew it would change the game. These gzipped XML files are about 150 to 250 MB. Unzipped, about 1 GB each.
Better, these files are based on Europe PMC identifiers, hopefully resolving the issue with using dates in the queries.</p>

<p>Now, parsing a 1 GB XML files is a total non-issue. I have done it plenty of times before. Just use a
<a href="https://en.wikipedia.org/wiki/Simple_API_for_XML">Simple API for XML</a> (SAX) parser. This is a streaming parser
giving you full control of how to parse things. It is ideal for this siutation: you just keep the current
paragraph of text in memory and release that when done with that paragraph. That is, you do not have to read
the full file in memory, just the bits you are interested in. I used this for my Chemical Markup Language
patches for Jmol and JChemPaint back in the nineties.</p>

<p>Last weekend I finally made the jump. Use SAX to extract the <code class="language-plaintext highlighter-rouge">&lt;p&gt;</code> elements one by one, running OSCAR on
them, filter with OPSIN, output that name, and clear the memory. Effectively, each gzipped file processes
with a Groovy script in about 1 to 2 hours.</p>

<p>The output is a mesmerizing stream of scientific literature (which I will use until someone points me to a Java
CLI library that creates a Matrix-style falling letters equivalent), tho less so as a static image:</p>

<p><img src="/assets/images/jats_analysis.png" alt="" /></p>

<p>In this plot, an <code class="language-plaintext highlighter-rouge">x</code> means a new article to be processed. Each <code class="language-plaintext highlighter-rouge">.</code> and <code class="language-plaintext highlighter-rouge">o</code> that follows is a single <code class="language-plaintext highlighter-rouge">&lt;p&gt;</code>
element and the difference is that an <code class="language-plaintext highlighter-rouge">o</code> means at least one IUPAC name was detected in the paragraph.</p>

<p>Each gzipped file gives 400 to 500 new IUPAC names. Indeed, going from 288 thousand to 300 thousand
was a matter of a day and a half. And earlier this afternoon we passed the 400 thousand IUPAC names.
With about 230 gzipped files. Now, I am going back in time, and the sizes of these files are shrinking:
Another 500 files and the size has dropped to around 125 MB, so a rough estimate suggests that
we will end up with 650 to 700 thousand names this way. This will be completed in a few weeks (and mostly
because I need to focus first on other things again, because I can use our computing cluster do this).</p>

<p>Regarding the original goal, fortunately, we are still publishing at a higher rate every year, and
more and more articles are available as open access. So, I still have good hopes we will reach the
<em>1 million IUPAC names</em>. Also, keep in mind, we know how to boost this by simple name variations to
several millions, even with the <a href="https://codeberg.org/BlueObelisk/iupac-names/commit/30ddfd96c3ec6e6a5840be0ada1bdbd40972490e">400 thousand</a>
we have today.</p>

<p>Oh, and <a href="https://github.com/BlueObelisk/iupac-names/issues/4">our next milestone</a> will be in the pocket
before I visit <a href="https://cheminf.uni-jena.de/">Christoph Steinbeck’s cheminformatics team</a> in Jena!</p>

      <h4>References</h4>
      <ul>
      
      
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="iupac"/><category term="textmining"/><category term="xml"/><category term="europepmc"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/05/02/one-million-iupac-names-5-a-new-approach.html">
      <![CDATA[ About fifteen months ago a new project started: One Million IUPAC names: ]]>
    </summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chem-bla-ics.linkedchemistry.info/assets/images/jats_analysis.png"/>
    <media:content xmlns:media="http://search.yahoo.com/mrss/" medium="image" url="https://chem-bla-ics.linkedchemistry.info/assets/images/jats_analysis.png"/></entry>
  
  <entry>
    <title type="html">Open Science Festival Limburg</title>
    <link href="https://chem-bla-ics.linkedchemistry.info/2026/04/27/open-science-festival-limburg.html" rel="alternate" type="text/html" title="Open Science Festival Limburg"/>
    <published>2026-04-27T00:00:00+00:00</published>
    <updated>2026-04-27T00:00:00+00:00</updated>
    <id>https://doi.org/10.59350/0qjrj-59n30</id>
    <content type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/04/27/open-science-festival-limburg.html">
      <![CDATA[ <p>One of the things I have been busy with in the past weeks (besides contributing to grant proposals) is the organization
of the June 11-12 <a href="https://www.openscience-maastricht.nl/events/open-science-festival-2026/">Open Science Festival Limburg</a>!
Open Science Festivals provide a great platform to talk science in an open way, with people that all find reuse of
science more important. The creativity present at such events is just so energizing.</p>

<p>This is the second time a local Open Science Festival is organized in the area, with an
<a href="https://www.openscience-maastricht.nl/events/open-science-festival-2023/">Open Science Festival Maastricht</a> in 2023.
Of course, a year later Maastricht also hosted the <a href="https://opensciencefestival.nl/open-science-festival-2024">2024 national Open Science Festival</a>.</p>

<p>This year’s event is co-organized by the (new) Open Science Community Parkstad and the
<a href="https://opensciencefestival.nl/open-science-festival-2024">Open Science Community Maastricht</a>,
with interdisciplinarity as main theme. The event will take day over two days, <a href="https://www.openscience-maastricht.nl/11-june-2026/">June 11 in Heerlen</a>,
and <a href="https://www.openscience-maastricht.nl/12-june-2026/">June 12 in Maastricht</a>. You can <a href="https://www.openscience-maastricht.nl/events/open-science-festival-2026/">register</a>
for either or for both days. Each days have plenary sessions and several parallel workshops,
where you will dive into one of the many aspects of open science.</p>

<p>Looking forward to welcoming you in Limburg!</p>

      <h4>References</h4>
      <ul>
      
      </ul>
      ]]>
    </content>
    
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="openscience"/><category term="osflimburg"/>
    
    <summary type="html" xml:base="https://chem-bla-ics.linkedchemistry.info/2026/04/27/open-science-festival-limburg.html">
      <![CDATA[ One of the things I have been busy with in the past weeks (besides contributing to grant proposals) is the organization of the June 11-12 Open Science Festival Limburg! Open Science Festivals provide a great platform to talk science in an open way, with people that all find reuse of science more important. The creativity present at such events is just so energizing. ]]>
    </summary></entry>
  
</feed>
