{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2026/05/02/one-million-iupac-names-5-a-new-approach.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/gqtbx-jta57",
      "url": "https://chem-bla-ics.linkedchemistry.info/2026/05/02/one-million-iupac-names-5-a-new-approach.html",
      "title": "One Million IUPAC names #5: a new approach and 400k names",
      "content_html": "<p>About fifteen months ago a new project started: <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html\">One Million IUPAC names</a>:</p>\n\n<blockquote>\n  <p>Thus, the idea came up, can we create a set of 1 million unique IUPAC names found in literature?</p>\n</blockquote>\n\n<p>We started out with using <a href=\"https://europepmc.org/\">Europe PMC</a> to get JATS XML files for the full texts of open access articles.\nParsing the XML is easy and the text paragraphs are passed through OSCAR and OPSIN. That has not changed.</p>\n\n<p>What did change last weekend is something I had long on my todo list (but life interfered). The first approach\nwas to ask for named entities using the Europe PMC APIs. But I quickly realized that with OSCAR and OPSIN we could\nget more names out of the articles. The next step was to move from Google Colab to a command line script.\nThat gave another boost, as explained in <a href=\"http://localhost:4000/2025/04/27/one-million-iupac-names-2-the-100-thousand-milestone.html\">this second post in the series</a>.\nWe reached 200 thousand names in <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/06/09/one-million-iupac-names.html\">june 2025</a>\nbut then things slowed down again in the growth. <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.html\">Two months later</a>\nwe only had 75 thousand more. However, plenty of discussion was happening and there turned out to be\nother, larger collections of IUPAC names under an open license. Millions of names, actually.</p>\n\n<p>But another problem emerged. We were still using the Europe PMC API and were basically asking for open access\narticles between two dates. Practically, the API could answer requests between 1 and max 3 days. Beyond that,\ntimes outs and 404s became an issue. Moreover, because these dates are publications dates and not the dates\non which the JATS were deposited, I had to got back to previous months and redo the queries. That gave another\n5 thousand names since last August. Something had to change.</p>\n\n<h2 id=\"the-new-approach\">The new Approach</h2>\n\n<p>Europe PMC, however, also provides the JATS XML files as download on <a href=\"https://europepmc.org/ftp/oa/\">their FTP site</a>.\nAlready that <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.html\">august 2025</a> I had\na prototype and knew it would change the game. These gzipped XML files are about 150 to 250 MB. Unzipped, about 1 GB each.\nBetter, these files are based on Europe PMC identifiers, hopefully resolving the issue with using dates in the queries.</p>\n\n<p>Now, parsing a 1 GB XML files is a total non-issue. I have done it plenty of times before. Just use a\n<a href=\"https://en.wikipedia.org/wiki/Simple_API_for_XML\">Simple API for XML</a> (SAX) parser. This is a streaming parser\ngiving you full control of how to parse things. It is ideal for this siutation: you just keep the current\nparagraph of text in memory and release that when done with that paragraph. That is, you do not have to read\nthe full file in memory, just the bits you are interested in. I used this for my Chemical Markup Language\npatches for Jmol and JChemPaint back in the nineties.</p>\n\n<p>Last weekend I finally made the jump. Use SAX to extract the <code class=\"language-plaintext highlighter-rouge\">&lt;p&gt;</code> elements one by one, running OSCAR on\nthem, filter with OPSIN, output that name, and clear the memory. Effectively, each gzipped file processes\nwith a Groovy script in about 1 to 2 hours.</p>\n\n<p>The output is a mesmerizing stream of scientific literature (which I will use until someone points me to a Java\nCLI library that creates a Matrix-style falling letters equivalent), tho less so as a static image:</p>\n\n<p><img src=\"/assets/images/jats_analysis.png\" alt=\"\" /></p>\n\n<p>In this plot, an <code class=\"language-plaintext highlighter-rouge\">x</code> means a new article to be processed. Each <code class=\"language-plaintext highlighter-rouge\">.</code> and <code class=\"language-plaintext highlighter-rouge\">o</code> that follows is a single <code class=\"language-plaintext highlighter-rouge\">&lt;p&gt;</code>\nelement and the difference is that an <code class=\"language-plaintext highlighter-rouge\">o</code> means at least one IUPAC name was detected in the paragraph.</p>\n\n<p>Each gzipped file gives 400 to 500 new IUPAC names. Indeed, going from 288 thousand to 300 thousand\nwas a matter of a day and a half. And earlier this afternoon we passed the 400 thousand IUPAC names.\nWith about 230 gzipped files. Now, I am going back in time, and the sizes of these files are shrinking:\nAnother 500 files and the size has dropped to around 125 MB, so a rough estimate suggests that\nwe will end up with 650 to 700 thousand names this way. This will be completed in a few weeks (and mostly\nbecause I need to focus first on other things again, because I can use our computing cluster do this).</p>\n\n<p>Regarding the original goal, fortunately, we are still publishing at a higher rate every year, and\nmore and more articles are available as open access. So, I still have good hopes we will reach the\n<em>1 million IUPAC names</em>. Also, keep in mind, we know how to boost this by simple name variations to\nseveral millions, even with the <a href=\"https://codeberg.org/BlueObelisk/iupac-names/commit/30ddfd96c3ec6e6a5840be0ada1bdbd40972490e\">400 thousand</a>\nwe have today.</p>\n\n<p>Oh, and <a href=\"https://github.com/BlueObelisk/iupac-names/issues/4\">our next milestone</a> will be in the pocket\nbefore I visit <a href=\"https://cheminf.uni-jena.de/\">Christoph Steinbeck’s cheminformatics team</a> in Jena!</p>",
      "summary": "About fifteen months ago a new project started: One Million IUPAC names:",
      "image": "https://chem-bla-ics.linkedchemistry.info/assets/images/jats_analysis.png",
      "date_published": "2026-05-02T00:00:00+00:00",
      "date_modified": "2026-05-02T00:00:00+00:00",
      "tags": ["iupac","textmining","xml","europepmc"],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
