{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "chem-bla-ics",
  "description": "Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.",
  "home_page_url": "https://chem-bla-ics.linkedchemistry.info/",
  "feed_url": "https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.json",
  "icon": "https://chem-bla-ics.linkedchemistry.info/assets/images/chem-bla-ics_logo.png",
  "language": "en",
  "authors": [
    {
      "name": "Egon Willighagen",
      "url": "https://orcid.org/0000-0001-7542-0286",
      "_orcid": "0000-0001-7542-0286"
    }
  ],
  "items": [

    {
      "id": "https://doi.org/10.59350/krw9n-dv417",
      "url": "https://chem-bla-ics.linkedchemistry.info/2025/08/09/one-million-iupac-names-4.html",
      "title": "One Million IUPAC names #4: a lot is happening",
      "content_html": "<p>A lot is happening. If you have been following this project more closesly, you may have already seen some interesting updates, but\nI will post it here too. First, a quick recap. In March I started a new <a href=\"http://blueobelisk.org/\">Blue Obelisk</a> project to\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html\">collect CCZero IUPAC names</a>\nfrom primary literature (paper still pending). It turned out we can automate that, while legally not violating any laws or licenses.\nIn April I reported on <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/04/27/one-million-iupac-names-2-the-100-thousand-milestone.html\">some tweaks</a>\nboosting the efficiency of the use of the API. I also reported on some possible further steps, including how to use the extracted\nnames to create a larger set. Indeed, in June I could <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/06/09/one-million-iupac-names.html\">report to have passed the 200k IUPAC names</a>,\nwhich with the idea from April gave us more than 1M IUPAC names.</p>\n\n<p>In this post I want to give an update.</p>\n\n<h2 id=\"275k-iupac-names\">275k IUPAC names</h2>\n\n<p>I have continued running the scripts to detect new IUPAC names in full text, open access papers in <a href=\"https://europepmc.org/\">Europe PMC</a>,\nbut something more awesome actually did much more since the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/06/09/one-million-iupac-names.html\">June post</a>:\nin July I received a <a href=\"https://github.com/BlueObelisk/iupac-names/pull/13\">pull request</a> from <a href=\"https://github.com/mnietfeld\">mnietfeld</a>\nwith more than 40 thousand unique and new IUPAC names from the <a href=\"https://www.beilstein-journals.org/bjoc/\">Beilstein Journal of Organic Chemistry</a>\n(see also <a href=\"https://www.linkedin.com/posts/beilstein-institut_openaccess-bjoc-fair-activity-7351596602660167681-0Z0r/\">their LinkedIn post</a> or\n<a href=\"https://archive.is/DZOnP\">this archived version</a> that doesn’t require an account).\nWhile Europe PMC provides these articles too (and actually one of the first I analyzed), a lot of these names come from supplementary\ninformation, not provided by Europe PMC. Thanks!</p>\n\n<p>This is focusing on names from primary literature, but there is more happening. Because I want to restrict the above project to\nnames from primary literature (and supplementary information is still that), I have not been sure what to do with other collections\nyet, and they have been coming in. I have been <a href=\"https://github.com/BlueObelisk/iupac-names/issues?q=is%3Aissue%20label%3Aother\">taking notes</a>\nin the project issue tracker, for future reference (like now, here). I have not forgotten about these!</p>\n\n<h2 id=\"other-large-collections-of-iupac-names\">Other large collections of IUPAC names</h2>\n\n<p><strong>4M, CCZero</strong><br />\nLet’s start with the news yesterday. The <a href=\"https://www.ebi.ac.uk/about/teams/chemical-biology-services/\">Chemical Biology Services team</a>\n<a href=\"https://chembl.blogspot.com/2025/08/unleashing-4-million-iupac-names-into.html\">released 4 million IUPAC names from patent literature as CCZero</a>!\nThe CCZero license/waiver makes it compatible with our list. Their Zenodo release:</p>\n\n<blockquote>\n  <p>… contains IUPAC names text-mined from patents (US, WIPO, EPO, Chinese, Japanese).</p>\n</blockquote>\n\n<p>The post also includes a nice example of the complexity of IUPAC names which makes the counting of unique names tricky:\n<code class=\"language-plaintext highlighter-rouge\">O-methylphenol</code> and <code class=\"language-plaintext highlighter-rouge\">o-methylphenol</code>. Thanks, Noel and the rest of the EMBL-EBI team!</p>\n\n<p><strong>2.3 million, CC-BY</strong><br />\nAnd then <a href=\"https://github.com/haydn-jones\">Haydn Jones</a> was one of the earliest <a href=\"https://github.com/BlueObelisk/iupac-names/issues/9\">to coin in</a>,\nand <a href=\"https://doi.org/10.5281/zenodo.15077270\">released 2.3 million IUPAC names</a> under the CC-BY license.</p>\n\n<p><strong>850k, CCZero</strong><br />\nWikidata also turnes out to have many IUPAC names. <a href=\"https://github.com/Adafede/\">Adriano</a> found more than 850 thousand IUPAC\nnames, see <a href=\"https://github.com/Adafede/wd-labels-to-iupac\">this project</a>.</p>\n\n<p>Next week I will do some comparisons of the datasets with a clear Creative Commons license.</p>\n\n<h2 id=\"even-more\">Even more</h2>\n\n<p>Beyond these five data releases, there is more. PubChem and other databses have millions of names, but often these are\ngenerated by proprietary software. These IUPAC name collections may be under some license agreement, and thus not compatible\nwith Open Science. This is why it is so important that we very clearly know where these names are coming from.</p>\n\n<p><strong>5-6 million, license unclear</strong><br />\nI also learned about <a href=\"https://chempile.lamalab.org/\">ChemPile</a> about which <a href=\"https://www.linkedin.com/in/adrian-mirza-chem/\">Adrian Mirza</a>\nexplained me it has <a href=\"https://www.linkedin.com/feed/update/urn:li:activity:7330626142611062784\">about 5-6 million IUPAC names</a>.\nBut the source of this list of names is not yet clear to me.</p>\n\n<p><strong>Names from PhD theses and preprints</strong><br />\nI also want to give a shout out to <a href=\"https://github.com/BlueObelisk/iupac-names/issues/15\">Peter Murray-Rust</a>s proposal\nto start extracting IUPAC names from PhD theses. There have been projects to extract chemistry from PhD thesis in the\npast, and this will yield a lot of unique names. Please ping Peter, if you want to get involved in his idea!</p>\n\n<h2 id=\"whats-next\">What’s next</h2>\n\n<p>I am so excited with all these efforts and very grateful with the contribution by Beilstein. I really hope more Open Science\npublishers will follow, like perhaps the Royal Society of Chemistry for which it should be easy, with their\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2007/02/01/rsc-first-publisher-to-go-semantic.html\">Project Prospect</a> background!</p>\n\n<p>I am also excited by the release by ChEMBL under CCZero. That will allow the <a href=\"https://www.wikidata.org/wiki/Wikidata:WikiProject_Chemistry\">WikiProject Chemistry</a>\nuse this for Wikidata!</p>\n\n<p>So, I have one week left to write the article about the work we started in March. The outlook is bright. I played last\nweek with the Europe PMC full text downloads and can confirm that should yield thousands of additional names from the\nfull texts. A single download file gave me more than two thousand new unique names. I think the 500k IUPAC names\nis absolutely in reach with purely the full texts from Europe PMC.</p>\n\n<p>This brings us to the end of 2025. By then, we should have a many millions of openly-licensed IUPAC names.\nAnd by March 2026, I hope we reached the 1M IUPAC names extracted from primary literature. That will require some\ncreativity and enthusiasm, but sounds feasible!</p>\n\n<h4>References</h4>\n<div class=\"csl-bib-body\">\n    <div class=\"csl-entry\">Jones, H. (2025). <i>OPSIN PubMed Names</i> [Dataset]. Zenodo. https://doi.org/10.5281/ZENODO.15077270 <b>[cito:citesAsRecommendedReading]</b> <a href=\"https://doi.org/10.5281/zenodo.15077270\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.5281/zenodo.15077270\">Scholia</a></div>\n    <div class=\"csl-entry\">O’Boyle, N., &#38; Bosc, N. (2025). <i>IUPAC names text-mined from patents by SureChEMBL</i> [Dataset]. Zenodo. https://doi.org/10.5281/ZENODO.16755947 <b>[cito:citesAsRecommendedReading]</b> <a href=\"https://doi.org/10.5281/zenodo.16755947\">CrossRef</a> <a href=\"https://qlever.scholia.wiki/doi/10.5281/zenodo.16755947\">Scholia</a></div>\n  </div>",
      "summary": "A lot is happening. If you have been following this project more closesly, you may have already seen some interesting updates, but I will post it here too. First, a quick recap. In March I started a new Blue Obelisk project to collect CCZero IUPAC names from primary literature (paper still pending). It turned out we can automate that, while legally not violating any laws or licenses. In April I reported on some tweaks boosting the efficiency of the use of the API. I also reported on some possible further steps, including how to use the extracted names to create a larger set. Indeed, in June I could report to have passed the 200k IUPAC names, which with the idea from April gave us more than 1M IUPAC names.",
      
      "date_published": "2025-08-09T00:00:00+00:00",
      "date_modified": "2025-09-03T00:00:00+00:00",
      "tags": ["iupac","beilstein","chembl","europepmc"],
      "_references": [
        
          
          
            { "url": "https://doi.org/10.5281/zenodo.16755947", "doi": "10.5281/zenodo.16755947"
            , "cito":
              
              
                [ 
                  "citesAsRecommendedReading"
                  
                 ]
              
             }
            ,
          
        
          
          
            { "url": "https://doi.org/10.5281/zenodo.15077270", "doi": "10.5281/zenodo.15077270"
            , "cito":
              
              
                [ 
                  "citesAsRecommendedReading"
                  
                 ]
              
             }
            
          
        ],
      
      
      
      
      
      
        "authors": [ { "name": "Egon Willighagen", "url": "https://orcid.org/0000-0001-7542-0286" } ]
      
    }

  ]
}
